Beginner· Foundations· 8 min read

How a large language model produces an answer

From training data to tokens to a reply, the process behind the chat window, step by step.

Dr. Michael D’Rosario
Host & Editor · September 29, 2026
↗ in ✉

When you ask a large language model a question, it can feel as though something familiar is happening behind the screen. You ask a question, the model considers what it knows, develops an answer and then writes that answer back to you.

That is an understandable interpretation because it resembles what another person might do. It is also a poor description of how a large language model actually produces its response.

A language model does not ordinarily begin with a completed answer sitting somewhere inside the system waiting to be expressed. Instead, it constructs the response progressively. Given the text and other information available in its context, the model calculates which pieces of language are plausible next steps, selects one, adds it to the context, and then performs the process again. What eventually appears to us as a coherent paragraph, argument or explanation is produced through many successive predictions.

Understanding that process is one of the most useful pieces of AI literacy because it helps explain several characteristics of generative AI that otherwise seem contradictory. It helps explain how a model can produce an insightful analysis without having retrieved it from a database, why changing a few words in a prompt can alter the answer, why the model can sometimes correct itself halfway through a response, and why something that writes with extraordinary confidence can nevertheless produce information that is wrong.

It starts with your context

Suppose you ask a language model:

Why did inflation increase after the pandemic?

The model does not simply search internally for a stored answer labelled “post-pandemic inflation”. Your question becomes part of what is called the context, which is the information currently available to the model when producing its response.

Depending on the AI system you are using, that context may contain considerably more than the sentence you have just typed. It might include earlier messages in the conversation, documents you have supplied, instructions governing how the model should behave, information retrieved from another source, or the results of tools the system has used.

The model’s task is then conditioned on that context. In simplified terms, it is asking: given everything available to me at this point, what should come next?

That final phrase, “at this point”, is important because the context changes continuously while the answer is being generated.

The model does not quite read words

Before the model can process your question, the text is divided into units called tokens.

Tokens are not necessarily whole words. A common word might correspond to a single token, while an unusual or complicated word might be represented by several. Punctuation, numbers and fragments of words can also be tokens. The exact division depends on the tokenizer used by the particular model.

The sentence:

Inflation increased rapidly.

might therefore be represented not simply as four words, but as a sequence of numerical token identifiers corresponding to pieces of that text.

This is an important conceptual shift. The model does not directly receive language in the form that we experience it. Text is converted into numerical representations that a neural network can process.

Those token representations are then mapped into high-dimensional mathematical representations known as embeddings. Rather than treating every token as an isolated symbol, embeddings allow the model to represent complex relationships between language and concepts mathematically.

Words and concepts that occur in related contexts can acquire related representations, although the actual structure is considerably richer than simply placing similar words next to one another on a map. Modern language models represent relationships across many dimensions simultaneously, allowing information about meaning, syntax, context and other patterns to interact.

Context changes meaning

Language is inherently contextual. Consider the word “bank” in these two sentences:

The bank increased its lending rate.

We sat on the bank of the river.

The sequence of letters is identical, but the meaning is different. A useful language model therefore cannot simply attach one fixed meaning to the word “bank”. It has to interpret that token in relation to the surrounding tokens.

This is where the transformer architecture that sits behind most contemporary large language models becomes particularly important.

A transformer uses a mechanism known as attention to calculate relationships between tokens in the context. In simplified terms, attention allows the model to determine which other parts of the available information are particularly relevant when processing a given part.

If the text concerns interest rates, mortgages and monetary policy, those relationships help establish that “bank” probably refers to a financial institution. If the surrounding text refers to water, fishing and a river, a different interpretation becomes appropriate.

The same principle operates across much longer and more complicated contexts. A model analysing a report may connect a claim on one page with a qualification appearing much earlier. In a conversation, it may interpret “that result” by relating the phrase to something discussed several messages ago. In programming, it may connect a variable being used in one section of code with the place where it was defined.

Attention is one of the mechanisms that allows the model to construct these contextual relationships.

Then comes the prediction

After processing the context through many layers of the neural network, the model produces a probability distribution over possible next tokens.

Imagine, in an extremely simplified example, that the model has begun an answer with:

Inflation increased after the pandemic partly because…

There are many possible continuations. The model might assign different probabilities to tokens corresponding to “demand”, “supply”, “households”, “energy”, “governments” and thousands of other possibilities.

It is not simply choosing between five words. Modern language models typically operate with vocabularies containing many thousands of possible tokens, and the model assigns some probability to each possible continuation.

A simplified distribution might look something like this:

demand 22%
supply 18%
household 11%
energy 8%
government 6%
other possibilities 35%

These numbers are illustrative, but the principle is important. The model does not necessarily identify one uniquely correct next word. It calculates a distribution of possibilities given the context.

The system then uses that distribution to determine which token comes next. Exactly how it does this depends partly on the generation settings. Some systems behave relatively deterministically, favouring the highest-probability continuation, while others permit more variation by sampling from a wider range of plausible possibilities.

This is one reason the same prompt can sometimes produce different answers.

And then it does it again

Suppose the selected token corresponds to “demand”.

The model now has a new context:

Inflation increased after the pandemic partly because demand…

It processes that updated sequence and generates another probability distribution for what should follow.

Perhaps the next token produces “recovered”.

The context becomes:

Inflation increased after the pandemic partly because demand recovered…

The model calculates again.

Then perhaps:

Inflation increased after the pandemic partly because demand recovered faster…

And again:

Inflation increased after the pandemic partly because demand recovered faster than…

This process continues until the response is complete.

The important point is that the model is not simply revealing a paragraph that already existed internally. Every generated token becomes part of the information used to generate the next token.

The answer is therefore path dependent. An early choice affects the probability of later choices, which influence subsequent choices in turn. By the time the model has generated several paragraphs, its own preceding response has become a substantial part of the context shaping what follows.

This helps explain why prompts matter so much. Changing the framing at the beginning changes the initial context, which changes the early probability distributions, which can send the generated response down a substantially different path.

But “predicting the next token” is only part of the explanation

This is where descriptions of language models can become misleading.

It is technically reasonable to say that an LLM generates language by predicting tokens, but the phrase can create the impression that the system is performing something comparable to the predictive text function on a mobile phone.

The crucial difference is what produces those predictions.

A large language model has been trained across enormous quantities of data to adjust billions of numerical parameters. Through that training, the model develops complex internal representations of relationships within language and the information represented through language.

Producing a useful next-token prediction for a complicated question may therefore require the model to represent a considerable amount of structure.

If I write:

The Reserve Bank increased interest rates because…

predicting a sensible continuation requires relationships between concepts such as inflation, monetary policy, demand, expectations, borrowing costs and central banking.

If I instead ask the model to explain whether a particular regression specification adequately addresses endogeneity, generating a useful answer requires very different relationships.

The output is still generated token by token, but the probability distribution behind each token is being produced by a neural network containing an enormous quantity of learned structure.

Saying that an LLM “just predicts the next token” therefore confuses the immediate generation mechanism with the capability of the system producing the prediction.

Where did all of those relationships come from?

They came principally from training.

During pre-training, the model is repeatedly presented with sequences of data and trained to predict parts of those sequences. When its prediction differs from the training target, the parameters within the neural network are adjusted slightly to reduce the error.

This happens repeatedly across a vast training process.

The result is not a database containing copies of everything the model has seen. Instead, training alters the numerical parameters of the neural network so that the model becomes progressively better at representing and predicting relationships found within the training data.

This distinction helps explain why asking an LLM “Where did you learn that?” can be problematic.

There may be no single document for the model to retrieve. A particular output might reflect statistical relationships learned from many examples distributed throughout its training data. Unless the AI system has access to a search or retrieval mechanism, it may not have a reliable means of connecting a generated claim to its original source.

Why can it get facts wrong?

Once we understand the generation process, hallucinations become less mysterious.

Suppose you ask a model for the title of a particular academic paper. If the model has strong enough representations associated with the paper, it may produce the correct title. If its representation is weak or ambiguous, however, it still has to generate a continuation.

A plausible academic title has statistical structure. So does an author’s name, a journal citation, a DOI and a publication year.

The model can therefore generate something that fits those patterns even when the underlying reference does not exist.

This is not the model consulting a database, failing to find the answer and consciously deciding to invent one. It is the generation process continuing under conditions where plausibility and truth have come apart.

The distinction becomes particularly important because language models can be extremely good at plausibility.

Why does asking it to “think” sometimes help?

More capable AI systems can also allocate additional computation to a problem before producing a final response, and models can be trained to perform intermediate reasoning operations that improve performance on tasks requiring several steps.

The underlying principle is still related to sequential generation. Breaking a problem into intermediate steps can change the context available for subsequent computation, allowing earlier results to inform later ones.

Consider a mathematical problem requiring several calculations. Attempting to jump directly from the question to the final number places considerable pressure on a single prediction. Producing intermediate representations gives the model additional structure from which subsequent steps can be generated.

Something similar can occur in analysis. Identifying assumptions, separating evidence from inference, comparing competing explanations and then constructing a conclusion provides a different computational path from immediately generating an answer.

This does not make the model infallible, nor does a lengthy reasoning process prove that the reasoning is correct. It does help explain why additional computational effort and structured intermediate processing can improve performance on some tasks.

The model is only one part of the AI system

There is one final complication because many products that we casually call an “LLM” now contain much more than the language model itself.

When you ask a contemporary AI assistant a question, the surrounding system might decide that it needs current information, search the internet, retrieve several sources, pass relevant material back to the model, allow the model to interpret those sources, call another software tool to perform a calculation, and then use the resulting information to construct its final response.

What appears to the user as a single conversation may therefore involve a sequence something like:

Question → context → model → tool selection → retrieval → new context → model → answer

In an organisational setting it can become more elaborate:

Question → model → company database → retrieved records → model → analytical tool → result → model → proposed action → approval → business system

This distinction matters because the capabilities and reliability of the resulting AI system no longer depend solely on the language model. They depend on the quality of retrieval, the available tools, the information supplied to the model, the permissions attached to those tools and the checks applied before an output becomes an action.

Understanding how a large language model produces an answer therefore gives us a better way to interpret what appears on the screen. The response is not simply retrieved knowledge, nor is it a completed thought being translated into words. It is a computationally generated sequence in which each step is conditioned on an evolving context, using relationships represented through training to determine what should come next.

That mechanism is simultaneously more mechanical and more capable than our conversational interface suggests. Once we understand it, several apparent contradictions begin to disappear. An LLM can produce genuinely useful analysis and still invent a reference. It can represent sophisticated relationships while making a basic factual error. It can generate different answers to the same question without either answer having been stored anywhere beforehand.

The useful question is therefore not simply whether an AI “knows the answer”. For anyone trying to use these systems well, the more important questions are what information the model had available, what process produced the response, whether the task requires generation or retrieval, and what evidence we have that the resulting answer is correct.

Listen to the episode
What is the Intelligent Economy

More AI Literacy

Beginner · 6 min read

What AI actually is, and what it isn’t

By Dr. Michael D’Rosario

Advanced · 10 min read

AI literacy as public infrastructure

By Dr. Michael D’Rosario

Advanced · 12 min read

AI regulation explained: from principles to practice

By Dr. Michael D’Rosario