Menu

OpenAI scientist Noam Brown: The true upper limit of AI may not be able to measure

As large language models gradually tackle complex tasks such as reasoning, automated research, and cybersecurity, traditional methods of evaluating models are facing new challenges.


For a long time, the release of models has been accompanied by a report of results consisting of various benchmark tests in areas such as mathematics, programming, scientific question answering, network security, and knowledge reasoning, which are then compared horizontally with the previous generation of models.



OpenAI researcher Noam Brown recently pointed out in an article that when models use more reasoning steps, invoke more tools, or perform longer searches and tests when answering questions, a single score becomes increasingly difficult to accurately reflect the model’s actual capabilities.



The main points of Brown are as follows:The performance of large models depends not only on the models themselves but also on the amount of computing resources they receive during the inference phase.When evaluating models in the future, we cannot simply ask “How many points did the model get?”; we should also answer another question: How many tokens did the model consume, and at what cost and running time was this achievement obtained?


He suggests that the industry should shift from focusing on "individual achievements" to "the performance of the inference computation volume curve" from the very beginning, and regard the inference budget as an essential variable for evaluating model capabilities and as part of artificial intelligence security policies.


Traditional report cards may underestimate the performance gap between new models.


Brown used the market reaction after the release, as indicated by GPT-5.5, to illustrate the limitations of traditional model rankings.


According to his description,GPT-5.5In the early stages of the release, what attracted attention from the outside world were a set of benchmark test results that were not particularly impressive.GPT-5.4In comparison, the scores of the new model have improved, but from the perspective of traditional score tables, the increase seems to be limited. As a result, some users are taking a wait-and-see attitude or even expressing skepticism about the new version.


However, within a few hours of the model being made available, some users noticed that as developers and researchers began testing more complex tasks, GPT-5.5 exhibited more pronounced generational differences in long-chain reasoning, continuous execution, and handling of complex problems. Brown believes that this phenomenon, where "practical experience has significantly improved, but the list scores have changed only slightly," reflects the fact that traditional evaluation methods do not fully capture the capabilities of the model.


The problem is that the evaluation results of different models may not be based on the same reasoning budget.


In traditional evaluation frameworks, researchers often choose a set of test configurations for each model in order to maximize the results as much as possible, and then place the final scores in the same table. This may seem fair, but it can obscure a key variable: some models may continue to improve significantly after receiving more inference tokens, more calls, or longer running times; other models may reach their performance ceiling earlier.


The cybersecurity assessment cases presented by Brown show that if the so-called models are only compared based on the final results under the condition of “calculating the maximum amount during testing,”GPT-5.5ComparisonGPT-5.4The advantages may not be very prominent. However, if they are, the tokens will be used to control the number of instances, the cost of reasoning, or the latency at the same level, and then the performance of different models will be observed.GPT-5.5The improvement in capabilities will be even more noticeable.



In other words, the difference between the models is not only reflected in the final scores but also in the efficiency of using additional computational resources for inference.


Why can't it be simple? "Just keep running until the performance no longer improves?"


An intuitive solution is to continue adding additional computing resources to each model until its performance reaches a stable level (the “plateau phase”), and then compare the maximum capabilities of these models.


Brown believes that this idea may not be feasible in practice. The reason is that for the new generation of models, the period of performance plateau may be much later than expected, and it may even be difficult to observe even within the actually affordable budget range.


He cited the automated research experiment initiated by Andrej Karpathy as an example. In these experiments, the model continued to show improvement in performance even after numerous trials. Even after hundreds of experiments, the curve of improvement was not entirely flat (i.e., it continued to show a trend of improvement).



Brown also mentioned the cybersecurity assessment results of the AI Security Institute in the UK. In this assessment, Mythos was included, as well asGPT-5.5Some models, including the ones mentioned, continued to show improved task performance even after using more than 100 million tokens in total.



This phenomenon indicates that models can utilize longer running times and larger computational budgets to continuously explore, experiment with, and refine strategies in complex tasks. More powerful models not only start with a higher baseline performance but may also be better at converting additional computational resources into effective capabilities.


Brown speculates that as model capabilities improve, the task cycles during which the models can operate effectively will also become longer. In the past, people might have observed a stabilization in model performance within relatively limited budgets; in the future, the upper limits of performance may be pushed further. For certain tasks, the so-called “plateau period” may no longer be an easily measurable state.


Moving from a single score to a “Performance-Cost Curve”


In response to this change, Brown suggests that model publishing organizations should alter the way benchmark tests are presented.


Rather than just announcing a final score, it would be better to mark the amount of reasoning computation on the horizontal axis, display the task performance on the vertical axis, and plot a complete curve showing the performance changes. The horizontal axis could use indicators such as the number of tokens, the cost of reasoning, or the actual running time.


This method can address questions that are difficult to explain using traditional score sheets. For example, which model performs better with the same budget? How much faster does a model improve when the budget is increased tenfold? Is the model approaching its performance limit? How do the cost-effectiveness of different models change?


Similar methods have already been used in some benchmark tests. Brown mentioned that the ARC-AGI evaluation attempts to measure the relationship between model scores and running costs, rather than just publishing a single score.



Another feasible approach is to establish clear scheme tokens, cost, or time constraints for the evaluation, and to inform the model of the budget information in advance. This approach is similar to how humans take standardized tests: whether it's the SAT (Scholastic Assessment Test) for college admissions in the United States or the International Mathematical Olympiad, participants need to complete the tasks within a fixed time frame. Model capabilities can also be compared under unified constraints.


However, Brown also pointed out that the number of different indicators is limited.


Due to the different parturers used by various models, as well as the generation speeds and units, the quantities cannot be directly compared across models. The cost of tokens may also vary. The cost is influenced by hardware utilization, batch processing, and the engineering implementation. Since runtime is not a perfect indicator, it is not an ideal measure either. Techniques such as “multi-agent cooperation” or best-of-N can generate multiple candidate answers in parallel, which may not significantly increase the waiting time experienced by users, but it does significantly increase the total amount of computation.


Nevertheless, he believes that any of the above indicators provide more information than a single score that deviates from the reasoning budget.


The issue of budget allocation for research and development is being extended to include the evaluation of artificial intelligence security.


Brown's discussion is not limited to the list of models. He believes that the reasoning budget also has a direct impact on the security management of cutting-edge models.


Before releasing cutting-edge artificial intelligence models, R&D organizations typically assess the potential for misuse, such as cyber attacks, biological risks, and chemical hazards. If the model exceeds certain risk thresholds, the organization may need to postpone the release or implement mitigatory measures, such as adding access restrictions and monitoring mechanisms, before deploying it.


The question is: If the model's capabilities improve with the increase in the amount of reasoning computation, how much reasoning budget should be used for security assessment?


In fact, ordinary users may only invest a few dollars or tens of dollars in a task. However, an organization with sufficient funds, a professional team, or a national actor may be willing to invest much more resources than ordinary users for a single goal. If the evaluation agency tests the model only with a low budget, it may underestimate its risk tolerance under high-resource conditions.


Brown uses the controversy surrounding the release of Gemini 3 Deep Think as an example. He points out that the benchmark test results for Deep Think were significantly higher than those of previous models, but the complete system documentation was not provided at the time of release to assess the risk capabilities of this version. This approach has drawn criticism from some researchers in the field of artificial intelligence security.




However, there seems to be an underlying issue behind the Brown controversy: AI companies and security organizations have not yet developed a consistent method for evaluating the capabilities of models under different reasoning budgets.


He speculates that Deep Think may not be a completely independently trained new model, but rather a reasoning framework system based on other existing models. This system could improve the performance of complex tasks by calling the model multiple times, generating candidate answers in parallel, automatically checking answers, and making iterative adjustments.


If this assumption holds true, then in theory, Deep Think can not only implement some of the capabilities demonstrated by the platform itself. As long as external developers are willing to invest sufficient effort in reasoning, they can also combine multiple models to build similar workflows. The role of Deep Think is more to encapsulate complex reasoning processes that typically require professional development skills into a product form that can be easily used by ordinary users.


Therefore, Brown and others believe that the真正 important question is not whether the system card was released separately from the product, but whether the research and development organizations have fully tested the capabilities that the basic model could have achieved when it was first released, under different inference budgets and framework strategies.


High-budget evaluations are difficult to implement comprehensively, but attempts can be made to make extrapolations.


Theoretically, an entity with sufficient resources might be able to invest more than $10 million in reasoning costs for a single task. However, security assessments typically involve thousands or even millions of tests. If an extremely high budget is used for each operation, the assessment cost would quickly become unfeasible.


Brown suggests that tests can be conducted within a relatively controllable budget for inference, and then performance under higher budget conditions can be improved based on the trend of model capabilities changing with the amount of computation. At the same time, the evaluation agency should clearly indicate the range of predictions and the uncertainties, rather than treating the computation results as definitive conclusions.



This method is similar to estimating the trends of change in larger systems using local data. It cannot replace actual testing, but it can help research and development organizations as well as regulatory agencies understand how the risk boundaries may change when models are given more time, tools, and computing resources.


However, Brown also acknowledges that long-term tasks may still pose problems that are difficult to solve with short-term experiments.


For example, if researchers want to determine whether an independent intelligent entity will exhibit target deviation, strategic deception, or other misaligned behaviors after running continuously for a year, the most reliable method might still be to let the intelligent entity operate for a sufficient length of time. Based on the results of just a few hours or days of experimentation, it may not be possible to detect the key changes in its long-term behavior.


This will create a new contradiction in reality: the development and release cycle of artificial intelligence models may only take a few months, while the task cycle for which intelligent bodies can operate continuously may become longer and longer. In the future, research and development institutions may face a special situation where the next generation of models will be nearing release before they have completed their maximum operating cycle.


Three suggestions: Make the reasoning budget a fundamental variable in model evaluation


In response to the issues mentioned in the capability assessment and security governance, Brown has put forward three specific recommendations.


Firstly, when releasing new models, artificial intelligence research and development institutions should publish benchmark test results under various inference budget conditions. Ideally, companies should provide performance curves with the number of tokens, cost, or runtime as the horizontal axis. At the very least, companies need to explain how much inference resources were actually used to achieve the single-point results.


Secondly, benchmark test rankings should record the consumption of reasoning resources, or set unified tokens, cost, or time limits for reference models. Currently, some evaluations have incorporated relevant variables, but a standard practice has not yet been established in the industry.


Thirdly, the computational resources for the reasoning phases of the Preparedness Framework for artificial intelligence companies and the Responsible Scaling Policy (RSP) should be carefully considered. When an organization determines whether a model has crossed a certain security threshold, it is necessary to evaluate not only the performance under a single configuration but also to assess multiple reasoning budget levels, and to make predictions about the risk capacity under higher budget conditions with a degree of uncertainty.


The industry has already recognized this issue, but the evaluation systems have not yet fully kept up with it.


Adding computational resources during the inference phase can improve model performance, and this is not a new discovery.


Since OpenAI released the o1推理 model in September 2024, the industry has generally believed that this model employs more reasoning steps when answering questions, resulting in better performance in mathematical tasks, coding, and complex analytical tasks. Research focusing on "the scalability of calculations during testing" or "the scalability of computational reasoning" has also gradually become an important direction in the development of large models.


However, Brown believes that even nearly two years after this trend emerged, the release of many advanced models still relies primarily on a single benchmark score for dissemination and comparison. Some security organizations can also re-examine the limits of model capabilities after using reasoning budgets that are dozens or even hundreds of times larger in their scaffolding systems.


As models become increasingly adept at using long-term running processes, multiple rounds of trial and error, and large-scale reasoning resources, the performance of traditional list-based approaches (such as explanation systems) is likely to continue to decline. Under various conditions, such as low-budget question-answering systems, high-budget deep research projects, multi-intelligence collaboration, and automated tool interactions, the same basic model may exhibit significantly different levels of capability.


Brown argues that in the future assessment of artificial intelligence capabilities, the reasoning budget should no longer be considered merely as auxiliary information during the testing process, but rather as a core parameter in the evaluation report, alongside other factors such as model size, training data, and context window.


From a broader perspective, this also means that the artificial intelligence industry is gradually moving away from the stage of defining a model with just one number. For capability assessment, product comparison, and security management, the truly important question may not be just what a model can do, but rather what it can achieve when it is given sufficient time, funding, and computing resources.


Reference link: https://x.com/polynoamial/status/2064210146558136827


The author comes from "Machine Heart" and the "Machine Heart Editorial Department".

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR