Advanced· Data & evidence· 12 min read

Evaluating AI systems: benchmarks and their limits

What test scores measure, what they miss, and how to evaluate a system for your own use.

Dr. Michael D’Rosario
Host & Editor · September 29, 2026
↗ in ✉

Every major AI model release now arrives accompanied by a collection of numbers. One model scores higher on mathematics, another improves on coding, another reaches a new level on a reasoning benchmark, while a smaller model claims comparable performance at a fraction of the cost. The numbers provide an appealing sense of order to a technology that can otherwise be difficult to compare, particularly for organisations trying to distinguish meaningful improvements in capability from the increasingly dense stream of claims surrounding each new generation of models.

Benchmarks are useful precisely because they create common tests. If several models answer the same questions under reasonably consistent conditions, their performance can be compared in ways that are more informative than asking each provider to demonstrate the tasks on which its own system performs particularly well. Benchmarks have played an important role in measuring progress in artificial intelligence, identifying weaknesses and providing some common basis for comparison across systems that may differ substantially in architecture, size and intended use.

The problem begins when a benchmark score is treated as a general measure of intelligence, quality or suitability. A model that performs better on a benchmark has demonstrated that it performs better on that benchmark, under the particular conditions in which the test was conducted. Whether this means it is the better model for analysing an organisation’s documents, writing its software, answering customer enquiries or operating an AI agent is a separate empirical question.

Understanding that distinction is central to evaluating AI systems properly.

A benchmark is a measurement instrument

The easiest way to understand a benchmark is to treat it like any other measurement instrument. It attempts to take something broad and difficult to observe directly, such as mathematical reasoning, coding capability, factual knowledge or instruction following, and measure it through performance on a defined set of tasks.

This immediately raises a familiar question from economics and statistics: does the measure adequately represent the underlying capability we actually care about?

Suppose we want to evaluate whether an AI system would make a useful research assistant. We could test whether it answers factual questions correctly, identifies information in documents, summarises research, interprets statistical results, distinguishes strong evidence from weak evidence and provides accurate citations. Each test would tell us something relevant, but none individually would establish that the model will perform well across the full range of activities involved in research.

The same issue appears whenever complex concepts are represented through indicators. An examination score can provide useful evidence about student knowledge without being identical to knowledge itself, while GDP measures a substantial component of economic activity without representing everything we might care about in economic welfare. Customer satisfaction scores contain information about customer experience, but they are not customer experience in its entirety.

Benchmarks should be interpreted in the same way. They are proxies for capabilities we care about, and their usefulness depends on how closely the thing being measured corresponds to the thing we actually need to know.

What exactly is being tested?

Consider a model reported to achieve 90 per cent accuracy on a reasoning benchmark. Before interpreting the number, we need to know what the benchmark contains, how the questions are structured and what the model was permitted to do while answering them.

Are the questions multiple choice or open ended? Do they test mathematics, science, logic or general knowledge? Were they designed for humans or specifically to challenge AI systems? Can the model use tools, execute code or retrieve external information? How much computation can it use before answering, and is it given one attempt or several?

These details matter because performance is conditional on the task. A model that performs exceptionally well on multiple-choice scientific questions may still struggle to interpret an ambiguous business problem in which the relevant information is distributed across several documents, important variables are missing and no single correct answer exists.

This does not make the benchmark invalid; it defines the domain within which its result should be interpreted. If a benchmark measures coding, it tells us something about coding, while a benchmark built around graduate-level scientific questions tells us something about performance on those questions. Problems arise when evidence from a relatively narrow domain is converted into the much broader claim that one model is simply more intelligent than another, without specifying the kinds of capability for which the comparison holds.

Benchmark performance is conditional on the test environment

Model performance is not produced by the model alone, because the same underlying model can produce substantially different results depending on the prompt, available context, reasoning budget, tools, sampling settings and number of attempts it is allowed to make. A system permitted to execute code or search external sources is solving a different problem from one required to answer entirely from information contained within the model and its immediate context.

This becomes particularly important as AI products move beyond simple model calls.

Imagine two systems completing a data analysis task. One uses a highly capable language model without additional tools, while the other uses a smaller model that can inspect the dataset, execute statistical software, check its calculations and revise its answer when a test fails. Asking which underlying model is better may tell us much less than asking which complete system produces the more reliable analysis at an acceptable cost.

The same issue becomes even more pronounced with agents. An agent’s performance depends not only on whether its language model can reason about the problem, but on whether it selects the appropriate tools, supplies them with the correct information, interprets their results properly, recovers from failures and recognises when the task is complete.

As AI becomes more deeply integrated with software, evaluation consequently needs to move beyond model capability towards system performance. What matters operationally is not the intelligence of a component in isolation, but the reliability of the process in which that component is embedded.

There is also a problem of teaching to the test

Benchmarks become more complicated once they become important.

If a particular benchmark influences reputation, investment, procurement decisions or perceptions of technical leadership, model developers have strong incentives to improve performance on it. There is nothing inherently problematic about this. Improving the capabilities measured by a useful benchmark is precisely what a benchmark is intended to encourage, particularly when the measure remains closely aligned with capabilities that users value.

The difficulty arises when optimisation becomes increasingly specific to the measure itself.

This is a version of the problem commonly associated with Goodhart’s law: once a measure becomes an important target, behaviour begins to respond to the measure, which can gradually weaken the relationship between improvements in the metric and improvements in the underlying outcome.

Developers may train models on tasks resembling benchmark questions, refine systems around known weaknesses or construct inference strategies specifically designed to maximise performance under benchmark conditions. Scores can continue to rise even when the corresponding improvement in broader capability is smaller than the headline numbers suggest.

The problem is hardly unique to artificial intelligence. Schools can teach towards standardised tests, organisations can optimise performance indicators without materially improving the underlying service, and researchers can respond strategically to metrics used in academic evaluation. In each case, the measurement system begins to influence the behaviour it was originally intended only to observe.

The appropriate response is not to abandon measurement, but to interpret mature benchmarks with greater care and continue developing evaluations that remain informative as systems adapt to existing tests.

Has the model seen the test before?

Large language models are trained on enormous quantities of information, including substantial volumes of publicly available material, which creates a distinctive problem for benchmarks whose questions or close variants have appeared online.

If benchmark material is contained in training data, strong performance may partly reflect familiarity with the questions, answers or closely related examples rather than the general capability the benchmark is intended to measure. This problem, usually described as contamination, becomes more difficult as benchmarks become older, more prominent and more widely discussed.

Establishing contamination is not always straightforward because training datasets can be extraordinarily large and model providers do not necessarily publish complete details of their contents. Even when exact benchmark questions have been excluded, related solutions, discussions, explanatory material or derivative examples may have appeared elsewhere in the training corpus.

The problem becomes progressively more important for established benchmarks because successful tests tend to become widely reproduced. Questions and solutions appear in academic papers, repositories, tutorials, discussion forums and training materials, increasing the possibility that later models have encountered information closely related to the evaluation.

Newer benchmarks often respond by using private test sets, newly created questions or tasks whose answers change over time. None of these approaches completely resolves the problem, but they follow an important principle: a test provides stronger evidence of generalisation when there is good reason to believe that the model is solving a genuinely unseen problem rather than reproducing patterns associated with familiar material.

Benchmarks can run out of room

A benchmark is most informative when it distinguishes meaningfully between the systems being evaluated. Once nearly every leading model scores close to the top, small differences become difficult to interpret and the test tells us progressively less about capabilities beyond the threshold it was designed to measure.

This is benchmark saturation, and it is an inevitable consequence of technological improvement when the test itself remains fixed.

As models become more capable, evaluations that once separated systems clearly can become routine. Researchers then develop harder tests containing more complex questions, longer reasoning chains or more realistic tasks, which means that public discussion can appear to move continually from one benchmark to another as older measures lose their ability to discriminate among leading systems.

None of this means that the underlying progress is artificial. It means that measurement has to move with capability, just as an examination designed to distinguish among beginners becomes less useful when administered to a group of experts.

A 95 per cent score on a saturated benchmark and a 60 per cent score on a deliberately difficult new benchmark therefore cannot be compared as though they occupied a common scale. The percentage acquires meaning from the difficulty, construction and purpose of the test on which it was produced.

Average performance can conceal the failures that matter

Suppose an AI system correctly completes 95 per cent of tasks in an evaluation. At first glance, that sounds like excellent performance, but its operational significance depends heavily on what happens in the remaining 5 per cent.

If the failures are distributed randomly across low-consequence tasks, the system may be extremely useful. If they are concentrated in unusual cases that happen to involve the largest financial transactions, vulnerable customers, ambiguous instructions or security-sensitive operations, the same average accuracy can provide a dangerously reassuring picture of reliability.

This matters particularly for AI because organisations often want to deploy these systems precisely where the volume of work makes individual human review difficult. A model might perform extremely well across thousands of routine cases while failing systematically when documents become unusually long, information is contradictory, a customer uses an uncommon language or the correct response requires recognising that essential information is missing.

Operational evaluation therefore needs to examine not only how frequently a system fails, but how those failures are distributed and how visible they are when they occur. A model that is slightly less accurate on average but reliably signals uncertainty may be more useful than one with a higher headline score whose errors are difficult to distinguish from its correct answers.

Reliability is partly a property of accuracy, but it is also a property of failure. Organisations need to understand the circumstances in which a system becomes unreliable, the consequences of those failures and whether the surrounding process can detect them before they matter.

The cost of being wrong is not constant

Standard accuracy metrics frequently treat errors symmetrically, with a correct answer receiving credit and an incorrect answer being counted as an error regardless of its consequences. Real organisations rarely operate under conditions in which every error has the same economic or social cost.

If an AI system incorrectly categorises an internal document, the consequence may be negligible. If it incorrectly approves a substantial financial transaction, denies an eligible customer access to a service or introduces a security vulnerability into production software, the significance of the error is quite different.

Evaluation should therefore often weight performance according to consequence rather than treating every observation as equally important.

A customer service system might be permitted to answer routine administrative questions with considerable autonomy while cases involving financial hardship, formal complaints or account security receive different treatment. A financial system might automate thousands of low-value reconciliations while requiring stronger controls around a relatively small number of high-value transactions.

Under these conditions, the most useful metric may not be the percentage of answers that were correct, but the frequency, severity and concentration of consequential errors. That is a more demanding evaluation problem, although it is also much closer to the question organisations actually need to answer before delegating meaningful work to an AI system.

Better performance may come at a cost

Model comparisons frequently focus on capability while treating computational cost and response time as secondary considerations, yet both can become central once a system is deployed at scale.

Suppose Model A correctly completes 94 per cent of a particular task while Model B achieves 96 per cent. If Model B costs ten times as much to operate and takes substantially longer to respond, the additional two percentage points may represent excellent value for one application and poor value for another.

For a high-value analytical task performed a few hundred times each year, the additional cost may be trivial relative to the value of improved performance. For a consumer service handling tens of millions of interactions, small differences in cost per request can accumulate into substantial expenditure.

Latency behaves similarly. A model that takes several minutes to complete a difficult research task may be entirely acceptable if the alternative is several hours of human work, whereas the same delay in a real-time customer interaction could make the service impractical.

The relevant objective is therefore not maximum benchmark performance irrespective of cost. It is sufficient performance for the task, delivered with an appropriate combination of reliability, speed and resource use.

A smaller model can be the better model

This has an important practical implication because the most capable general-purpose model is not automatically the appropriate model for every task.

Suppose an organisation needs to classify incoming requests into eight well-defined categories. A smaller model that performs the classification accurately, responds quickly and costs very little may be preferable to a frontier model with substantially broader capabilities, because the larger model’s ability to solve difficult mathematics, generate sophisticated software or analyse complex research contributes little to the narrow objective being pursued.

The same reasoning applies to specialised and locally deployed models. An organisation may rationally prefer a somewhat less capable general model because it can operate within controlled infrastructure, meet stringent latency requirements, run without an external connection or process a very large volume of requests economically.

This is an ordinary optimisation problem rather than a contest to identify the most intelligent model. Additional capability has value when it contributes to the objective, and its economic value declines when the organisation is paying for capabilities that the task does not require.

Evaluation should consequently begin with the work rather than with the leaderboard.

Human preference is useful, but complicated

Some AI outputs do not have a single objectively correct answer. Which explanation is clearer, which summary is more useful, which piece of writing is stronger and which conversational response better satisfies the user’s request are questions that often require human evaluation.

People can compare outputs and express preferences, producing useful evidence about which systems users find more helpful. These evaluations matter because many of the tasks we want AI to perform involve qualities that cannot easily be reduced to factual accuracy or a predetermined answer key.

Human preference, however, introduces its own measurement problems. People may prefer longer answers even when they contain unnecessary information, reward confident language even when the underlying evidence is weak, or favour agreeable conclusions over more carefully qualified ones. Preferences can also vary considerably according to expertise, occupation, culture and the purpose for which the output will be used.

An answer preferred by a casual user may not be preferred by a specialist evaluating technical accuracy, while an answer that feels satisfying in a short interaction may prove less useful when incorporated into a professional workflow.

The evaluator consequently becomes part of the measurement system. A high score on human preference tells us something valuable about how a particular group of people responds to a model’s outputs, but it should not automatically be interpreted as evidence that those outputs are more accurate, more rigorous or better suited to every use.

The benchmark that matters most may be your own work

For organisations choosing between AI systems, public benchmarks provide a useful starting point because they can narrow the field, identify broad capabilities and reveal major weaknesses without requiring every prospective user to reproduce large technical evaluations independently.

They should rarely be the final stage of evaluation.

If an organisation intends to use AI to classify its customer enquiries, it should test models on representative customer enquiries. If the intended task is analysing contracts, evaluation should involve the kinds of contracts the organisation actually encounters. If AI will generate software within a particular technical environment, the test should resemble that environment rather than an abstract coding exercise selected because it is convenient to score.

This does not necessarily require an enormous evaluation programme. A carefully constructed collection of representative tasks can reveal a great deal, particularly when it includes not only common cases but also the difficult, ambiguous and consequential cases most likely to expose weaknesses.

Performance can then be assessed across dimensions that matter for the intended use, including factual accuracy, completeness, adherence to instructions, consistency, latency, cost and severity of errors. Different dimensions can be weighted according to the actual business problem rather than according to whatever happens to be prominent on a public leaderboard.

The resulting evaluation may attract considerably less attention than a global benchmark, but it will usually be much more useful for deciding whether a system belongs inside a particular workflow.

Test the system you intend to deploy

Organisation-specific evaluation also needs to reproduce the system that will actually be used.

If employees will interact with a model through a particular system prompt, retrieval system, document library and set of tools, that configuration should be tested. If an agent will be permitted to perform several actions before escalating to a person, the evaluation should assess that process rather than the language model in isolation. If a customer-facing system will operate with limited context and strict response-time requirements, those constraints belong inside the evaluation rather than being removed to give the model ideal conditions.

Otherwise, the organisation risks measuring a system that will never exist in production.

This is particularly relevant when interpreting demonstrations. A carefully engineered demonstration may use ideal prompts, selected examples, substantial computational resources and repeated attempts, whereas the operational system will encounter ambiguous requests, incomplete data, unusual edge cases and users who do not conveniently phrase their questions in the way the system was designed to receive them.

A useful evaluation should reproduce those conditions closely enough that success on the test provides meaningful evidence about likely performance in practice.

Measure improvement against the current process

AI evaluation also requires a counterfactual because performance has little economic meaning without an alternative against which it can be assessed.

A system achieving 90 per cent accuracy sounds impressive until we learn that the existing process achieves 98 per cent. Conversely, an AI system achieving 80 per cent may create substantial value if the current process is expensive, slow and only 65 per cent accurate, particularly if the remaining difficult cases can be identified and escalated.

The relevant comparison is therefore not perfection but the best realistic alternative.

That alternative may be a person performing the task manually, conventional software, a smaller model, an outsourced service or a hybrid process in which AI completes routine work and people handle exceptions. The comparison also needs to extend beyond accuracy because an AI system that produces similar quality while reducing turnaround from three days to ten minutes may create considerable value, while a system that marginally improves accuracy but requires extensive human review may create much less.

Performance needs to be evaluated as part of a production process, including the labour, infrastructure, review and error costs associated with using the system, rather than as an abstract property of the model.

Evaluation should continue after deployment

A final limitation of one-off benchmarks is that AI systems do not operate in static environments. The work changes, customer behaviour changes, documents are updated, models are replaced, retrieval systems accumulate new information and employees alter how they interact with the technology.

A system that performed well during initial testing may therefore perform differently six months later, even if the organisation has not consciously changed the workflow.

This makes evaluation an ongoing operational activity rather than a procurement exercise completed before deployment. Organisations can monitor samples of outputs, track error rates, compare automated decisions with later outcomes and examine whether performance differs systematically across groups or types of cases. Where the underlying model, data source or system configuration changes materially, previous evaluation results may need to be reconsidered rather than assumed to remain valid.

This discipline is already familiar in statistical modelling. A predictive model is not considered permanently reliable because it performed well on a test set at the time it was developed; its performance can deteriorate as relationships in the underlying data change. AI systems warrant the same treatment, particularly when they are being used repeatedly inside consequential organisational processes.

Read the leaderboard, then evaluate the work

Benchmarks remain valuable because, without them, claims about model capability would depend even more heavily on demonstrations, anecdotes and marketing. Well-designed benchmarks create common standards, reveal technical progress and provide evidence about capabilities that would otherwise be difficult to compare systematically.

Their limitation is not that they measure nothing, but that they measure something specific, under particular conditions, and the result becomes misleading when that specificity is forgotten.

When a new model reaches the top of a leaderboard, the useful questions concern what tasks produced the score, whether the test still distinguishes effectively among leading models, whether contamination is plausible, what tools and computational resources were permitted, how performance varies across categories and what kinds of errors sit behind the average.

After those questions comes the one that matters for actual adoption: can this system perform the work we need it to perform, under the conditions in which that work actually occurs, at an acceptable cost, with errors whose frequency and consequences we understand well enough to manage?

Public benchmarks can provide useful evidence towards that decision, but the strongest evaluation ultimately comes from testing the system against the work it is actually being considered to do.

Listen to the episode
What is the Intelligent Economy

More AI Literacy

Intermediate · 9 min read

Reading an AI claim like an economist

By Dr. Michael D’Rosario

Intermediate · 8 min read

Your data and AI: what happens to what you share

By Dr. Michael D’Rosario

Advanced · 10 min read

AI literacy as public infrastructure

By Dr. Michael D’Rosario