Artificial intelligence produces no shortage of claims about artificial intelligence. AI will increase productivity by 30 per cent. Automation will save employees ten hours a week. A new model is 40 per cent better than its predecessor. AI will eliminate millions of jobs, create millions more, reduce operating costs, increase revenue, democratise expertise and transform entire industries.
Some of these claims will prove broadly correct. Some are already supported by credible evidence. Others depend on assumptions that become considerably less impressive once they are made explicit.
Economists have a useful habit when confronted with claims of this kind. Before deciding whether a number is large, small, good, bad or even particularly meaningful, we tend to ask what exactly has been measured, compared with what, for whom, over what period and under what conditions.
That habit is especially valuable when reading claims about AI because the technology is changing quickly, the commercial incentives surrounding it are substantial, and impressive demonstrations are much easier to produce than evidence of sustained economic value.
The first question is not whether the claim sounds plausible. It is what would have to be true for the claim to mean what it appears to mean.
Compared with what?
Suppose a company reports that employees using an AI assistant completed a particular task 30 per cent faster.
That sounds like a productivity result, and it may be one. But before interpreting it, we need a counterfactual.
Thirty per cent faster than what?
Perhaps the comparison is with employees completing the same task without AI. Perhaps it is with a previous version of the software. Perhaps the comparison group received no training while the AI group did. Perhaps the task was deliberately selected because it was well suited to generative AI.
The counterfactual matters because an effect is always an effect relative to something else.
If an AI system allows an employee to complete a report in 70 minutes rather than 100, we can reasonably describe a 30 per cent reduction in task time. If another readily available software tool would have reduced the task to 75 minutes, however, most of the apparent AI effect is not an advantage relative to the relevant alternative.
This is the economic idea of opportunity cost applied to technology evaluation. The relevant comparison is rarely AI versus nothing. It is AI versus the best realistic alternative use of the same resources.
That alternative might be conventional automation, better software, process redesign, additional training, outsourcing, hiring another employee or simply leaving the process alone.
Without a meaningful counterfactual, an impressive improvement can tell us surprisingly little.
What exactly is the outcome?
“Productivity increased by 30 per cent” sounds precise while leaving almost everything important unspecified.
Was productivity measured as tasks completed per hour, revenue per employee, output per worker, time required for a standardised exercise or a subjective assessment by participants?
These are not interchangeable measures.
A programmer completing a coding exercise 30 per cent faster is evidence about the time required for that particular task under the conditions of the study. It does not automatically imply that software engineers using AI will produce 30 per cent more economically valuable software over a year.
The gap between those propositions contains a great deal of organisational reality.
Employees perform multiple tasks. Some tasks are complementary. Some create bottlenecks for others. Faster production can create additional review work. More code creates more code to maintain. More marketing material creates more material requiring approval. Time saved on one activity may be absorbed by meetings, administration or simply additional iterations of the same work.
A task-level efficiency gain is therefore not automatically a firm-level productivity gain.
When encountering an AI productivity number, ask which outcome was actually measured and how far the claim has travelled from that measure.
Who was studied?
Average effects can conceal important differences between people.
Suppose a study finds that generative AI improves average performance on a writing task by 15 per cent. The average may be accurate while telling us relatively little about who benefited.
Perhaps inexperienced workers improved by 30 per cent while highly experienced workers improved by 2 per cent. Perhaps lower-performing participants improved substantially while the strongest participants experienced little change. Perhaps some people became faster but less accurate.
These differences can matter more economically than the headline average.
If AI disproportionately raises the performance of less experienced workers, it may reduce the productivity gap between novices and experts for particular tasks. That could affect training, recruitment, wages and the organisation of work. If the gains instead accrue primarily to highly skilled workers who know how to use AI effectively, the technology may complement expertise and increase existing differences.
The same average effect can therefore imply very different labour-market consequences depending on its distribution.
Whenever a claim refers to “workers”, “employees”, “students”, “consumers” or “businesses”, ask who is actually represented in the evidence.
A study of several hundred consultants performing selected knowledge tasks is evidence about those participants and tasks. Whether the findings generalise to nurses, teachers, accountants, tradespeople or small business owners is a separate empirical question.
Who selected into the sample?
Selection is one of the quieter ways impressive AI results can be produced.
Consider a company reporting that teams using its AI tools are substantially more productive than teams that do not.
Perhaps the technology caused the difference.
But perhaps the employees who chose to use AI were already more technologically confident, more motivated or working in roles better suited to automation. Perhaps managers introduced the tool first in teams where they expected it to succeed. Perhaps unsuccessful users abandoned it and disappeared from the group later classified as AI users.
If adoption is not random, users and non-users may differ before the technology enters the picture.
This is a familiar identification problem. Observing that AI users perform better does not establish that AI caused the difference.
The stronger the causal language, the more closely we should inspect how the comparison was constructed.
Is it correlation or causation?
Generative AI is particularly good at producing causal stories because causal stories make sense of information.
Suppose an organisation introduces an AI assistant in January and customer response times improve by March. It is tempting to conclude that AI caused the improvement.
Perhaps it did.
But perhaps staffing increased in February. Perhaps seasonal demand declined. Perhaps the organisation simultaneously changed its workflow. Perhaps employees improved through experience. Perhaps response times were already improving before AI was introduced.
The observation that one event followed another does not tell us which counterfactual outcome would have occurred without the intervention.
This is why economists devote so much attention to identification. The question is not simply whether two things moved together, but whether we have a credible basis for attributing the change in one to the other.
Randomised experiments can help. So can natural experiments, difference-in-differences designs, regression discontinuities, instrumental variables and other approaches, depending on the problem. None is automatically valid simply because it has a technical name. Each depends on assumptions that need to be defended.
For everyday reading, however, you do not need to run an econometric model. You can begin with a simpler question: what else could plausibly have produced the observed result?
If the answer is “quite a lot”, the causal claim should be treated accordingly.
What is the denominator?
Percentages have an extraordinary ability to sound meaningful without telling us very much.
An AI system might reduce an error rate by 50 per cent. If the original error rate was 20 per cent, that is potentially substantial. If it fell from 0.02 per cent to 0.01 per cent, the operational significance may be very different.
A company might report a 200 per cent increase in AI-related revenue when that activity previously represented a tiny share of its business. A model might produce a 40 per cent improvement on a benchmark where the difference in absolute performance is only a few percentage points.
Whenever a percentage appears, find the denominator.
The same principle applies to large absolute numbers. A claim that AI could save an economy billions of dollars sounds substantial, but the interpretation changes once the figure is expressed relative to total expenditure, GDP, the size of the affected sector or the cost of implementing the technology.
Numbers need scale before they acquire meaning.
Is the improvement statistically significant, economically significant, or both?
Statistical significance and economic importance answer different questions.
With a sufficiently large sample, a very small difference can be estimated precisely enough to be statistically distinguishable from zero. That does not necessarily make the difference economically important.
Conversely, an estimated effect may be economically substantial but measured imprecisely because the study has limited data.
When reading an AI study, the useful questions therefore include the size of the estimated effect, the uncertainty around it and whether that magnitude would actually matter in practice.
Suppose an AI intervention improves a quality score by 0.4 per cent with very high statistical confidence. That may demonstrate a genuine effect while having little practical value once implementation costs are considered.
The p-value cannot answer that question.
Where are the costs?
Benefits tend to be easier to advertise than costs.
An AI system that saves employees 10,000 hours a year appears to have created a substantial economic benefit. But those hours are not free of context.
The organisation may have paid licence fees, integration costs, consulting expenses and additional computing costs. Staff may have spent time learning the system. Managers may need to review outputs. Security and governance processes may need to be established. Existing systems may require modification. Errors may create remediation costs.
There may also be transition costs that disappear over time and recurring costs that do not.
A credible economic assessment therefore asks about net benefit rather than gross benefit.
If AI saves $1 million worth of labour but costs $900,000 to operate, the relevant number is not the million-dollar saving. If it costs $1.2 million during implementation but only $200,000 annually thereafter, a single-year assessment may be equally misleading in the opposite direction.
Time matters, as does the boundary placed around the analysis.
Is saved time actually saved?
This deserves particular attention because time savings are among the most common measures used in generative AI studies.
Suppose employees report that AI saves them five hours each week. Multiplying five hours by the number of employees and their hourly wage can produce a very large estimated benefit.
But what happens to those five hours?
If employees use them to produce additional valuable work, the productivity gain may be real. If the time allows an organisation to avoid additional hiring, there may be a measurable financial effect. If it reduces unpaid overtime or allows workers to perform their existing jobs with less pressure, that is also a real benefit, although a different one.
If the saved time is simply absorbed by additional low-value tasks, the financial effect may be much smaller than the headline estimate suggests.
Time saved is an intermediate outcome. Its economic value depends on what happens next.
This distinction becomes increasingly important as AI adoption spreads because organisations may find that generating work faster simply moves the constraint elsewhere. Faster drafting creates more material for review. Faster analysis creates more recommendations requiring decisions. Faster software development creates more applications requiring maintenance.
Productivity depends on the system, not only the task.
Who benefits?
An intervention can increase total economic value while distributing that value very unevenly.
Suppose AI increases worker productivity by 20 per cent. The next question is who receives the resulting surplus.
Workers might receive higher wages, shorter working hours or improved job quality. Firms might receive higher profits. Customers might receive lower prices or better services. Some workers might benefit while others lose bargaining power or employment opportunities.
There is no reason to assume that productivity gains will automatically flow to any particular group.
This is especially important in claims about AI and labour markets. Statements that AI will “benefit workers” or “replace workers” can conceal considerable variation across occupations, industries, skill levels, bargaining arrangements and time periods.
Economics encourages us to separate the creation of surplus from its distribution.
Both matter.
What happens when everyone adopts it?
An advantage available to one firm does not necessarily remain an advantage once every competitor has access to the same technology.
Suppose AI allows a marketing agency to produce campaigns at half the previous cost. Initially, that may increase margins. As competing agencies adopt similar tools, however, competition may push prices down. Some of the productivity gain then passes to customers rather than remaining with producers.
The same dynamic can occur in software, consulting, design, legal services and many other forms of knowledge work.
This is the difference between a private return and a broader market effect.
At the firm level, early adoption may provide an advantage. At the industry level, widespread adoption can change prices, expectations, output volumes and the structure of competition. What begins as additional profit can eventually become the minimum capability required to remain competitive.
A claim about what AI does for one organisation should not automatically be treated as a claim about what happens when the technology diffuses across an entire market.
What changes in equilibrium?
This is where economic reasoning becomes particularly useful because technological effects rarely stop with the first-order change.
If AI makes software cheaper to produce, we should not assume society will simply produce the same amount of software using fewer programmers. Lower costs may increase demand for software, create new applications and change which organisations can afford to build it.
If AI reduces the cost of producing advertising, firms may produce more advertising, increasing competition for a resource that has not expanded at the same rate: human attention.
If AI makes written analysis much cheaper, organisations may demand more analysis rather than simply employing fewer analysts.
These responses can offset, amplify or redirect the initial effect.
The first-order question asks what happens when AI makes an existing task cheaper.
The economic question asks what people and organisations do once that task becomes cheaper.
That second question is often considerably more interesting.
Who is making the claim?
Evidence should be evaluated on its merits, but incentives still matter.
A company selling an AI product has an obvious interest in demonstrating that the product works. A consultancy advising organisations on AI adoption benefits from a market in which adoption appears valuable. A researcher may have incentives to publish novel findings. A company announcing workforce reductions may have reasons to attribute them to technological transformation rather than broader cost cutting.
None of this makes the claims false.
It tells us where scrutiny is warranted.
Economists are accustomed to thinking about incentives because behaviour responds to them. Information production is no exception. The relevant response is not cynicism, but attention to research design, definitions, data and whether independent evidence produces similar findings.
The stronger the commercial or institutional incentive attached to a claim, the more useful it becomes to understand how the number was constructed.
Forecasts are not observations
Claims about the future deserve their own category.
AI could add trillions of dollars to global GDP. Millions of jobs could be automated. Particular occupations may experience dramatic productivity improvements. Entire categories of work may change.
These forecasts can be useful, but they are not measurements of events that have already occurred.
They usually depend on assumptions about adoption rates, technical capability, investment, substitution between labour and capital, creation of new tasks, regulatory responses, consumer behaviour and the speed with which organisations reorganise themselves around the technology.
Small changes in those assumptions can produce very different outcomes.
When reading a forecast, therefore, ask for the mechanism rather than only the number. What assumptions generate the estimate? Which variables drive the result? What happens under alternative assumptions? Is the model estimating technical potential, economically viable adoption or actual expected adoption?
A technically automatable task is not necessarily an economically sensible task to automate, and an economically attractive technology is not necessarily adopted immediately.
Possibility, profitability and diffusion are different things.
Read the claim backwards
A useful way to evaluate a striking AI claim is to work backwards from the headline.
Suppose you read:
AI increases worker productivity by 35 per cent.
Before deciding what you think about it, reconstruct the claim.
What does productivity mean here? Which workers? Performing which tasks? Compared with what? Over what period? Was adoption random? What was the absolute change? Was quality held constant? Were implementation and review costs included? Did all workers benefit equally? Does the result persist after people become familiar with the technology? Does the effect survive outside the experimental setting? What happens when the technology is used across the entire organisation rather than for one task?
You may ultimately conclude that the 35 per cent result is compelling.
But you now know what the number means.
That distinction is the point.
Reading an AI claim like an economist is not about finding reasons to reject every optimistic forecast or impressive result. Nor does it require treating every claim as the beginning of an econometric investigation. It means resisting the temptation to let a headline number do more intellectual work than the evidence permits.
Ask about the counterfactual. Find the denominator. Identify the population. Separate correlation from causation. Look for selection. Distinguish task-level efficiency from organisational productivity. Count costs as well as benefits. Ask who receives the gains. Consider how behaviour changes once the technology becomes widespread, and treat forecasts as conditional statements about a future that has not yet occurred.
Most importantly, ask what the evidence actually establishes before deciding what you think it implies.
That is good practice when reading any economic claim. AI simply gives us a great many new claims on which to practise.