Trained Onis this AI training on my prompts?

How this record is made

Open Terms Archive saves a copy of each AI company's terms every time the page changes. Most of those changes are noise. This site keeps only the ones where the sentences about training on your prompts changed, and nothing is published on one reading alone.

1,122document versions recorded by Open Terms Archive
907identical to the previous version once normalised
64distinct states of the training clause
41candidate changes sent to review
22published after review, 22 by the readers

Why most recorded changes are not changes

On 24 September 2026 Open Terms Archive recorded a new version of ChatGPT's privacy policy. This is the entire difference:

[Free⁠](https://chatgpt.com/plans/free/?openaicom-did=0e733353-b1b0-467b-94de-262521eef0203a4fe887-51ca-4d69-94b5-8c95a7d81d1e&openaicom_referred=true)

A tracking ID inside one link rotated. The link also carries an invisible word-joiner character after "Free". Nothing a reader would see changed. OpenAI's pages do this several times a day, and a monitor that alerts on every recorded change reports that OpenAI rewrote its privacy policy several times a day.

The pipeline

  1. Normalise. Each recorded version is converted to plain comparable text. Link targets are dropped and link text kept. Invisible characters are stripped, including the word joiner (U+2060), zero-width spaces and soft hyphens. Unicode is normalised to NFC, typographic quotes become plain quotes, whitespace is collapsed, and screen-reader labels such as "(opens in a new window)" are removed.
  2. Segment. Each document is split into paragraphs, and paragraphs are split again at bullet points and table cells. Vendors put training language inside large purpose tables, and an edit elsewhere in the table must not look like a change to the clause.
  3. Locate by anchor. For every tracked document a person chose short, stable phrases that identify the sentences of the training clause, such as "use your inputs and outputs to train". A segment belongs to the clause when it contains one of them. The anchors are listed below, so anyone can check what is and is not tracked.
  4. Hash and collapse. The located clause is hashed. Consecutive versions with the same hash collapse into one clause state, however many times the rest of the document changed.
  5. Emit candidates. Each change of state becomes a candidate event. If no anchor matches at all, that is also an event, never silence: the vendor may have moved the clause, or Open Terms Archive's capture may have broken. Events that coincide with Open Terms Archive changing its own capture rules are flagged.
  6. Review by three readers. Each candidate goes to three language models from three companies: Claude (Anthropic), GPT (OpenAI) and Gemini (Google). Each sees only the old and new text of the clause, never the others' answers, and classifies the change as a change of position, a scope change, a new disclosure, wording only, or not a real change. When all three agree it is one of the first three, the change is published under the most cautious of their three labels, so it is never called more than the least that all three saw, and a further check picks the most cautious of their one-line summaries. Each such change is marked "three independent AI readers agreed". When they disagree on whether the change is real at all, when any of them is unsure, when the archive changed how it captures the page, or when the clause could not be found, a person decides. Wording-only and capture changes are never published. Registry rows work the same way: the quote is checked mechanically against the latest capture, and the answer label is confirmed by the same three readers or by a person.

What the dates mean

A date on this site is the date Open Terms Archive first recorded the new text, not the date the vendor decided the change. When Open Terms Archive went a while without capturing a document, the change page gives the window between the last capture of the old text and the first capture of the new.

Known gaps

Tracked documents and anchors

DocumentVersionsClause statesAnchors
ChatGPT, Business Privacy Policy 172 4 “we do not train our models on your”, “we do not use your business data for training”, “train our models on any”, “used to train other models”, “isn't used for training our models”, “trains its models in two stages”
ChatGPT, Privacy Policy 89 5 “to train the models that power”, “data we have collected from the internet to train our models”, “when we train”, “uses content from our services to improve and train”, “they can improve over time”, “we may use your content to train our models”, “do not train on my content”, “chats from temporary chat”, “if you enable temporary chat”, “sora has separate controls”, “in our training datasets”, “can be used to improve and train our models”, “if training is enabled”, “openly accessible on the internet”
ChatGPT, Terms of Service 62 3 “if you do not want us to use your content to train our models”
Claude.ai, Commercial Terms 3 1 “may not train models on customer content”
Claude.ai, Privacy Policy 660 6 “use your inputs or outputs to train”, “use your inputs and outputs to train”, “develop our language models that power our services”, “train our ai models as permitted under applicable laws”, “when you submit feedback, we disassociate”, “to train our trust and safety”, “conduct research (including model training)”, “conduct research, including training our models”, “legitimate interests to train and improve our ai models”, “may obtain personal data from the following”, “user inputs and outputs submitted as feedback”
Claude.ai, Terms of Service 6 3 “our use of materials”, “regarding any materials”
Windsurf (Codeium), Terms of Service 8 7 “use of autocomplete user content to improve services”, “use of chat user content to improve services”, “may use customer data for model training”, “elect to place limits on the use of”
Cursor, Privacy Policy 12 2 “use inputs or suggestions to train”
Cursor, Terms of Service 5 3 “including the training of language models”, “will not use content to train”
DeepSeek, Privacy Policy 12 5 “by training and improving our technology”, “to train our models and provide”, “to train and improve our technology”, “for training our models or optimizing”, “foundation model training”
GitHub Copilot, Terms of Service 20 3 “may be used for ai model training”, “for the purpose of training, developing, and improving”, “as input to ai features”, “affiliates may use your inputs and outputs”
Grok, Commercial Terms 6 2 “de-identified data and data retention”
Grok, Privacy Policy 9 3 “to train our models, but only with”, “publicly available on the internet to train our models”, “object to our use of your information to train”, “google apps content”
Grok, Terms of Service 8 2 “develop, train, test, improve and operate”, “used for product development or model training”
Jasper, Terms of Service 7 3 “train the artificial intelligence models developed by jasper”, “enhancing artificial intelligence models”, “object to the use of your customer property”
Le Chat, Commercial Terms 1 1 “do not use your data to train our models”
Le Chat, Data Processor Agreement 4 1 “training its artificial intelligence models”
Le Chat, Privacy Policy 2 1 “to train our artificial intelligence models”, “subject to your opt-out”, “which allows you to object”, “input and output data for model training”
Le Chat, Terms of Service 7 1 “we do not use your data to train our artificial intelligence models”, “mistral ai may use your feedback”
Microsoft Copilot, Privacy Policy 20 5 “we may use your data to develop and train our ai models”, “use children's data for model training”, “we use conversation data to train”
Perplexity, Data Processor Agreement 3 1 “will not be used for training”
Perplexity, Developer Terms 3 1 “customer content to train”
Perplexity, Privacy Policy 3 1 “email service information to create, train”

Source and licence

The underlying document history is Open Terms Archive's genai-contrib-versions, by Open Terms Archive contributors, under ODC-By 1.0. Everything published here is under the same licence. The data page has the files.