Beginner· Risks & ethics· 6 min read

When to trust an AI answer, and when not to

A short framework for deciding how much weight to give what a model tells you.

Dr. Michael D’Rosario
Host & Editor · September 29, 2026
↗ in ✉

The wrong question to ask about artificial intelligence is whether it can be trusted.

We do not usually talk about other sources of information in such absolute terms. We trust a calculator to perform arithmetic, but not to interpret employment law. We may trust an experienced colleague’s account of a meeting they attended while still checking their estimate of next year’s revenue. A government statistical agency may be an authoritative source for the unemployment rate while offering no particular authority on how an individual business should respond to it.

Trust is conditional on the task, the source of the information, the consequences of error and our ability to verify the result. AI should be treated in much the same way.

This becomes complicated because generative AI is unusually good at producing the appearance of authority. A weak answer does not necessarily look weak. It can be fluent, detailed, internally coherent and expressed with exactly the same confidence as a correct one. The model does not reliably become hesitant as the quality of its information deteriorates, which removes one of the signals we normally use when deciding whether another person’s answer deserves confidence.

The practical question is therefore not whether an AI system is trustworthy in general, but what would justify trusting this particular answer.

Start with what the answer depends on

Consider three requests.

Rewrite this paragraph more clearly without changing its meaning.

Explain the difference between nominal and real GDP.

What was Australia’s labour productivity growth in 2025?

They may all be answered through the same interface, but they present very different problems.

In the first case, almost everything the model needs is contained in the material supplied by the user. There is comparatively little external factual content to establish, and the user can inspect the result directly to determine whether the meaning has been preserved.

The second question depends on established conceptual knowledge. There is a correct distinction between nominal and real GDP, but it is well represented in economics texts and unlikely to have changed since the model was trained. A capable model should generally be able to explain it reliably, although a student or researcher may still want to verify important details.

The third question is different again. It asks for a particular statistic for a particular period. A plausible number is of little value if it is not the actual number, and the model’s general knowledge of the Australian economy is not enough to establish that. What matters is whether it has access to an appropriate current source, what measure of labour productivity is being used, and whether the figure can be traced to the underlying data.

The interface is the same. The evidentiary requirements are not.

Trust answers that can be inspected

AI is at its easiest to trust when the quality of the output can be evaluated directly.

If you ask a model to improve the structure of a paragraph, convert notes into a table, suggest alternative headings, explain a piece of code or summarise a document you have supplied, much of the evidence required to assess the result is already in front of you.

This does not mean errors are impossible. A summary can omit an important qualification, code can contain a subtle bug, and an edited paragraph can inadvertently alter meaning. The important difference is that there is a reference point against which the output can be checked.

This gives us a useful principle: trust can be higher when verification is cheap.

If an AI converts a set of meeting notes into a list of actions, you can compare the actions with the notes. If it restructures a dataset, you can check row counts, totals and samples. If it extracts fields from an invoice, those fields can be compared with the original document.

In these situations, AI is not asking us to accept its authority. It is performing work whose result can be inspected against the source material.

Be more cautious when the model is supplying the facts

The risk changes when the model itself becomes the apparent source of the information.

Ask an AI system:

What does the research say about four-day working weeks?

and it may produce a sophisticated synthesis discussing productivity, employee wellbeing, retention and implementation. Some or all of it may be accurate, but the polished answer does not tell you whether the underlying studies exist, whether they support the claims being made, whether contrary findings have been omitted, or whether the evidence comes from randomised studies, observational research, pilot programmes or corporate surveys.

This is where fluent synthesis can create an evidentiary illusion. The answer resembles the end product of research without necessarily having gone through the process required to establish the evidence.

The distinction is not whether the model has knowledge. Models can represent an enormous amount of useful information. The issue is whether the provenance of a particular claim matters.

For casual background information, approximate knowledge may be entirely adequate. For an academic paper, investment decision, legal argument, policy submission or medical decision, it may not be.

As the importance of provenance increases, so should the requirement that the answer be grounded in identifiable sources.

Sources change the trust equation, but they do not settle it

An AI answer supported by sources is generally easier to evaluate than an unsupported one, but the presence of citations should not end the assessment.

The first question is whether the sources actually exist. The second is whether they are authoritative for the claim being made. The third is whether they say what the AI claims they say.

These are different tests.

A model can cite a genuine paper while overstating its findings. It can use a source describing one population to support a claim about another. It can turn an association into a causal claim, treat a preliminary estimate as settled evidence, or cite an authoritative institution while overlooking that the relevant page contains superseded information.

The source therefore provides a pathway to verification rather than a guarantee of correctness.

For important work, the useful habit is not simply to ask AI for sources, but to require claims to be traceable to them. If a particular number matters, find the number in the source. If a research finding is important to the argument, check the study. If an AI-generated summary of legislation affects a decision, inspect the relevant provision or obtain appropriate professional advice.

The higher the consequences of error, the shorter the acceptable distance between the claim and its evidence.

Current information requires current access

Some questions have answers that change.

Who is the chief executive of a company? What is the current cash rate? How much does a particular product cost? What legislation is currently before parliament? What are the latest unemployment figures? Is a particular software feature available?

These are poor candidates for relying solely on what a model learned during training because the correct answer is partly a function of time.

A language model may have a strong representation of the Reserve Bank of Australia, monetary policy and the historical cash rate without knowing the rate today. Worse, it may generate an entirely plausible outdated figure.

For current questions, the relevant distinction is therefore between a model answering from its trained knowledge and an AI system with access to current information through search, retrieval, databases or other tools.

Even then, recency should be visible. A current source from this morning and an article written three years ago may both appear in a search result. The fact that an AI system can access the internet does not mean every statement it generates reflects the latest available information.

When time matters, check the date.

Numbers deserve particular suspicion

Language models can be very useful for quantitative work, but there is an important difference between reasoning about numbers and establishing them.

Suppose you ask:

Roughly how many Australian businesses would be affected by this regulation?

The model may know enough about Australian business demographics to generate a plausible estimate. It may even construct a sensible calculation. But if the required figure depends on the number of firms in particular industries, employee thresholds, exemptions and the latest business counts, a plausible estimate is not a substitute for those data.

The danger is that fabricated quantitative information often looks more credible than fabricated prose. A statement such as “approximately 184,000 businesses” carries an appearance of precision that “a substantial number of businesses” does not, even though the additional digits may have no evidentiary basis.

Whenever a number matters, ask where it came from.

If the answer is a calculation, inspect the inputs and method. If it is a statistic, identify the source, definition and period. If it is an estimate, distinguish observed values from assumptions. If the model cannot establish these, treat the number as a proposition requiring verification rather than as a finding.

Plausibility is not evidence

This is perhaps the most important habit to develop when working with generative AI.

Models are extraordinarily capable of generating plausible explanations.

If employee turnover increased after a restructuring, AI can explain why. If sales declined following a price increase, it can explain why. If one region has higher unemployment than another, it can construct an economic account of the difference.

The problem is that several explanations may be plausible at the same time.

Sales might have declined because prices increased, because a competitor entered the market, because advertising expenditure fell, because customer preferences changed, or because the decline was already occurring before the price change. A language model can construct a persuasive narrative around any of these possibilities.

The ability to explain an observation is not the same as evidence that the explanation is correct.

This distinction is especially important in strategy, economics, policy and research, where the interesting questions are often causal. AI can be extremely useful for generating hypotheses, identifying possible mechanisms and suggesting what evidence would distinguish between competing explanations. It becomes much less reliable when plausible mechanisms are quietly converted into established causes.

A good response to “Why did this happen?” should sometimes be “Here are the explanations consistent with the information available, and here is what we would need to know to distinguish between them.”

Watch for answers that are suspiciously complete

Real information is often incomplete.

Datasets have missing observations. Research literatures contain disagreements. Historical records have gaps. Organisational documents leave questions unanswered. Forecasts depend on uncertain assumptions.

Generative AI, however, is designed to continue generating.

That creates a subtle problem when the task itself implies that a complete answer is expected. Ask for ten reasons and the model has an incentive to supply ten. Ask for the market size in every country and empty cells become a problem to be solved. Ask for a complete chronology and the resulting narrative may quietly bridge periods where the evidence is weak.

A trustworthy answer is sometimes incomplete because the evidence is incomplete.

Expressions such as “I cannot establish this from the material provided”, “the available evidence does not distinguish between these explanations” or “a current source is required for this figure” are therefore signs of useful restraint rather than model failure.

The ability to leave a gap can be more valuable than the ability to fill one.

Expertise changes what you can safely delegate

There is a paradox in using AI for specialist work. The people who can gain the most from it are often those who are best able to identify when it is wrong.

An experienced programmer may use AI to produce substantial quantities of code because they can inspect the architecture, run tests and recognise problematic assumptions. An economist can use it to discuss identification strategies because they can identify when correlation is being confused with causation. A lawyer can use it to organise arguments while recognising when a proposition requires checking against the relevant authority.

A novice may receive exactly the same output without possessing the knowledge required to detect the error.

This means the appropriate level of reliance is partly determined by the user’s capacity to evaluate the result. AI can increase the productivity of expertise because experts can delegate work while retaining judgement over the output. Where that judgement is absent, external verification becomes more important rather than less.

The apparent sophistication of the answer should never be used as a substitute for the expertise required to assess it.

Consequences should determine the verification burden

Not every AI error matters equally.

If you ask for five alternative titles for a presentation and one is poor, very little has been lost. If you ask for a recipe and the suggested cooking time seems questionable, it is easy to check. If you ask for an interpretation of a contractual obligation involving millions of dollars, the consequences are different.

A useful way to think about reliance is therefore to combine two questions:

How likely is an important error?

What happens if I fail to detect it?

Where both are low, extensive verification may cost more than it is worth. Where either becomes substantial, additional checks become rational.

This is why there can be no universal rule that every AI output must be independently verified. Such a requirement would eliminate much of the productivity benefit of using AI for low-risk work. Equally, treating every output as sufficiently reliable because modern models perform well on benchmarks would ignore the uneven nature of model errors and the very different costs associated with them.

Verification should be proportionate to risk.

A simple trust test

Before relying on an AI answer, five questions are usually more useful than asking whether the model itself is trustworthy.

Can I inspect the answer directly?
If the task involves transformation, summarisation, classification or analysis of material you possess, compare the output with the underlying material.

Does the answer depend on external facts?
If it does, determine whether those facts came from identifiable sources or from the model’s internal representations.

Does the answer depend on current information?
If time matters, establish that the system has access to sufficiently recent information and check when the source was produced.

Is the model making an inference or reporting evidence?
A plausible explanation should not quietly become a factual conclusion.

What is the cost of being wrong?
The greater the consequence, the stronger the case for independent verification, specialist review or a different method altogether.

These questions produce a much more useful approach to trust because they attach confidence to the circumstances in which an answer was produced.

The greatest mistake is probably at either extreme. Treating every AI answer as inherently unreliable ignores the fact that these systems can perform many tasks with considerable accuracy and can make verification easier by working directly from supplied evidence. Treating fluent output as inherently authoritative makes the opposite error, confusing the model’s ability to produce convincing language with its ability to establish that every proposition within that language is true.

The more mature position sits between them. Use AI confidently where the task is well specified, the evidence is available, the result can be inspected and the consequences of error are manageable. Increase scrutiny as the answer becomes more dependent on external facts, uncertain inference, precise numbers, current information or specialist judgement.

Trust, in other words, should not begin with how intelligent the answer sounds. It should begin with whether you can establish why the answer deserves to be believed.

Listen to the episode
What is the Intelligent Economy

More AI Literacy

Intermediate · 9 min read

Hallucinations, bias and how to check the work

By Dr. Michael D’Rosario

Advanced · 10 min read

AI literacy as public infrastructure

By Dr. Michael D’Rosario

Advanced · 12 min read

AI regulation explained: from principles to practice

By Dr. Michael D’Rosario