Intermediate· Data & evidence· 8 min read

Your data and AI: what happens to what you share

Privacy, retention and training: what to know before you paste something in.

Dr. Michael D’Rosario
Host & Editor · September 29, 2026
↗ in ✉

Every time we use generative AI, we give it information. Sometimes that information is trivial, such as a request for dinner ideas or help rewriting a sentence. At other times it can include business plans, meeting notes, financial information, research data, source code, contracts, customer records or documents containing information about other people.

The simplicity of the interaction can obscure what is happening. We type information into a box, receive an answer and move on, much as we might use a calculator or search engine. Yet an AI service is a computing system, often operated by another organisation, and the information supplied to it has to be processed somewhere. Depending on the service being used, it may also be retained for some period, incorporated into account history, passed to connected tools or governed by particular settings concerning how the provider can use it.

This does not mean that entering information into an AI system automatically makes it public, nor does it mean that every conversation becomes training data. It does mean that the familiar instruction to “never put sensitive information into AI” is too crude to be particularly useful. The better question is what information is being shared, with which system, under what terms, for what purpose and with what controls.

What happens when you type something into an AI system?

Suppose you paste a paragraph into a generative AI service and ask for a summary.

At the most basic level, the text has to be transmitted to the computing infrastructure running the service so that it can be processed. The system converts the text into tokens, incorporates those tokens into the context available to the model and generates a response based partly on the information you supplied.

Your text is therefore being processed by the AI system. That much is unavoidable if the model is going to work with it.

What happens beyond that is a separate question.

The provider may retain the conversation as part of your chat history. It may keep operational records for security, abuse prevention or service administration. Depending on the product and its terms, data may or may not be eligible for use in improving models. Business and enterprise services may operate under different data arrangements from consumer products. An organisation may also have negotiated contractual conditions governing retention, processing location, access and deletion.

There is consequently no universal answer to the question “Does AI keep my data?”

You need to know which AI service you are talking about.

Processing is not the same as training

One of the most persistent misunderstandings concerns the difference between a model processing information and a provider using that information to train or improve future models.

If you paste a document into an AI application and ask it to summarise the document, the system necessarily processes that information as part of the request. This does not, by itself, mean that the underlying model is being retrained on the document.

Training is a separate computational process in which model parameters are adjusted using training data. A deployed model responding to a prompt generally does not rewrite its underlying parameters every time somebody speaks to it.

This distinction is important because people sometimes imagine AI as continuously absorbing every conversation into its permanent knowledge. That is not an accurate description of ordinary model inference.

However, a provider may separately collect eligible interactions and subsequently use them in processes intended to improve its systems, subject to the particular product, policies, contractual terms and user settings. Whether that occurs cannot be determined simply from the fact that the service uses a large language model.

“Did the model process this?” and “Can this information subsequently be used for model improvement?” are different questions.

Memory is different again

The introduction of memory and personalisation features adds another layer.

A system may retain information from previous interactions so that future conversations can be better adapted to the user. If you have told an AI assistant about your preferred writing style, current projects or recurring activities, a memory system may make some of that information available in later interactions.

This can feel as though the model itself has permanently learned about you, but the architecture may be quite different. Information can be stored outside the model and supplied to it when relevant, rather than encoded into the model’s underlying parameters.

The distinction is similar to the difference between changing a person’s brain and giving them a notebook to consult.

Conversation history, memory, retrieved documents and model training can all affect what an AI system appears to know, but they are not the same mechanism.

That distinction matters when considering privacy because deleting a conversation, changing a memory setting and excluding information from model training are conceptually different actions. The controls available for each depend on the particular service.

AI does not automatically make your information public

Another common concern is that information entered into an AI system will subsequently be given to another user.

That is not how a conventional language model works.

If you provide a confidential business plan to a model, another person cannot ordinarily ask the model to open your conversation and hand them the document simply because the model processed it. Their interaction has a different context, and they do not automatically gain access to yours.

This should not be confused with saying that no privacy or security risk exists. Data can be retained by service providers, accounts can be compromised, systems can be misconfigured, software can contain vulnerabilities, and connected applications can introduce additional pathways through which information moves.

There is also a broader issue concerning training data and whether models can reproduce information encountered during training. Models generally represent distributed statistical relationships rather than functioning as searchable copies of their training corpus, but memorisation of particular material can occur, especially for information that is repeated or distinctive.

The practical conclusion is neither that everything entered into AI becomes public nor that confidentiality can simply be assumed. The appropriate protections depend on the system and the sensitivity of the information.

Consumer and enterprise AI are not necessarily the same thing

This distinction is particularly important at work.

An employee opening a publicly available consumer AI application and pasting a customer document into it may be using a very different data environment from an employee accessing an enterprise AI service supplied and governed by their organisation.

An enterprise arrangement may provide contractual commitments concerning data use, administrative controls, access management, retention, auditability and integration with existing security systems. It may also allow the organisation to determine which models and services employees can access and what information those systems are permitted to process.

This means organisational policies that simply state “do not use AI with company data” can quickly become inadequate. The relevant question is often which information can be used with which approved systems.

A company might reasonably permit staff to use an approved enterprise AI environment with internal documents while prohibiting those same documents from being entered into an unapproved consumer service.

The data are identical. The governance environment is not.

The model may not be the only place your information goes

Contemporary AI systems increasingly use tools.

An AI assistant might search the internet, access cloud storage, query a customer relationship management system, read email, execute code or call an external service through an application programming interface. These capabilities make AI considerably more useful, but they also complicate the data question.

Suppose you upload a customer spreadsheet and ask an AI system to enrich the records using an external service. Information may now move between more than one system.

Similarly, an AI agent given access to email and a CRM may retrieve information from both, combine it in its working context and take an action in another application.

At this point, asking only how the language model handles data is insufficient. You need to understand the broader information flow.

What systems can the AI access? What information is sent to each tool? What is retained? Which organisation operates the service? What permissions has the AI been given? Can it write information back to those systems as well as read from them?

As AI systems become more agentic, data governance increasingly becomes a question about the entire chain of tools and services rather than about a single model.

Uploading a document creates another distinction

When you upload a document to an AI system, there are several ways the application might work with it.

For a relatively short document, some or all of the text may be placed directly into the model’s context. For a larger collection of documents, the system might extract the text, divide it into sections, create numerical representations of those sections and store them in a retrieval system. When you ask a question, relevant sections can then be retrieved and supplied to the model.

The latter approach is commonly associated with retrieval-augmented generation, or RAG.

This means the model does not necessarily have the entire document available during every interaction. The surrounding system may determine which parts appear relevant and provide those to the model.

For organisations building their own AI systems, this architecture creates important design choices. Documents can potentially remain within controlled infrastructure while only selected information is provided to a model. Local models can sometimes be used where information should not leave a particular environment. Access controls can restrict which employees or agents are able to retrieve particular records.

The question therefore shifts from “Can we use AI with confidential information?” to the much more useful question “What architecture and controls would allow us to use AI with this information appropriately?”

De-identification helps, but it is not magic

Removing names from data is often a sensible first step when working with sensitive information, but a dataset does not become anonymous simply because obvious identifiers have been deleted.

A record describing a person’s occupation, suburb, age, employer and unusual circumstances may still make that person identifiable, particularly when combined with information available elsewhere.

This matters for research, health, human resources and customer data, where combinations of apparently innocuous variables can reveal identity.

Good de-identification therefore considers whether a person could reasonably be re-identified from the remaining information rather than merely checking whether their name has been removed.

For many AI tasks, the model may not need identifiable information at all. If the objective is to classify themes in customer complaints, names, email addresses and account numbers may be irrelevant. Removing unnecessary fields reduces exposure without reducing the usefulness of the analysis.

This leads to a broader principle that applies well beyond AI: do not provide data that the task does not require.

Be particularly careful with other people’s information

People naturally think about privacy in terms of their own information, but workplace AI frequently involves information about someone else.

A manager may paste an employee’s performance review into an AI system. A consultant may upload a client’s documents. A researcher may work with participant responses. A business may process customer complaints containing names, contact details and personal circumstances.

The person operating the AI system may be comfortable sharing the information, but that does not mean they have the authority to do so.

This is where AI use intersects with existing obligations around privacy, confidentiality, research ethics, professional duties, contractual restrictions and information security. AI does not create all of these obligations, although it can create new ways of breaching them.

Before sharing information, the useful question is therefore not simply “Am I comfortable putting this into AI?” but “Do I have the authority to provide this information to this particular service for this purpose?”

A simple way to think before you share

Before entering information into an AI system, four questions can resolve much of the uncertainty.

What am I sharing?

Distinguish between public information, ordinary internal material, commercially sensitive information, personal information and information subject to specific legal, contractual or professional restrictions.

Which system am I sharing it with?

Identify whether you are using a consumer product, an organisation-approved enterprise service, an API, a locally operated model or another system entirely. The fact that two products use the same underlying model does not mean they have the same data arrangements.

What will happen to the information?

Check relevant retention, training, privacy and access arrangements rather than assuming them from the product name. If the AI is using external tools, consider those services as part of the information flow.

Does the system need this information?

If the task can be completed with fewer fields, de-identified information, synthetic examples or a smaller section of a document, provide only what is required.

These questions are considerably more useful than a blanket rule that sensitive information should never interact with AI because organisations already use externally operated computing infrastructure for email, cloud storage, payroll, customer management, analytics and many other functions. The relevant issue has always been whether the system, contractual arrangements and controls are appropriate for the information being processed.

AI should be treated with the same seriousness, while recognising that its capacity to combine, interpret and act on unstructured information creates some distinctive risks.

The important distinction is that “AI has my data” can describe several very different situations. Your information might exist temporarily in a model’s working context, remain in your conversation history, be stored in an external retrieval system, become available through a memory feature, be passed to another application, or be eligible for later use in improving a service. Those possibilities should not be treated as interchangeable.

Once they are separated, the practical questions become much clearer.

Know what you are sharing, know which system you are using, understand where the information can go, and provide no more than the task requires. The technology may be new, but the principle is familiar: access to information should follow purpose, authority and need, rather than convenience.

Listen to the episode
What is the Intelligent Economy

More AI Literacy

Advanced · 12 min read

Evaluating AI systems: benchmarks and their limits

By Dr. Michael D’Rosario

Intermediate · 9 min read

Reading an AI claim like an economist

By Dr. Michael D’Rosario

Advanced · 10 min read

AI literacy as public infrastructure

By Dr. Michael D’Rosario