Skip to main content
Back to Blog
6 min readBrass-SEO Team

Is Your Site in the Data That Trained ChatGPT?

There is a good chance your website already helped train an AI model. Not through a deal you signed. Through a free, open archive of the web that has been running since 2007.

That archive is Common Crawl, a nonprofit that crawls the public web and hands the results to anyone. In the paper that introduced GPT-3, filtered Common Crawl made up 60% of the training data. If your pages are public and you never blocked its crawler, they were eligible to be part of that corpus. This is a setting you can check and change, not a mystery you have to accept.

Quick Navigation


What Common Crawl Is

Brass-SEO points to Common Crawl as the clearest link between your robots.txt and AI. Common Crawl is a 501(c)(3) nonprofit that has published a free, open crawl of the web since 2007, adds 3 to 5 billion pages every month, and now spans more than 300 billion pages across fifteen years.

The data is free to download and has been cited in over 10,000 research papers. Its crawler is called CCBot. Because the corpus is large, open, and already cleaned up, it became the obvious raw material for teams building language models. You did not have to submit your site anywhere. If it was on the open web, it was in scope.

How It Became AI Training Data

Brass-SEO cites the GPT-3 paper for the scale of it. In Brown et al. (2020), filtered Common Crawl supplied 60% of GPT-3's training mix and 410 billion tokens, the single largest source in the model, drawn from 45 terabytes of raw crawl reduced to 570 gigabytes after filtering.

That was one model. Common Crawl remains one of the most common starting points for LLM training data, either used directly or cleaned into derivative datasets that other models learn from. The pattern holds: a public page that allows the crawler is a candidate for the next model's training set. The reach of your writing now extends past human readers to the systems those readers ask for answers.

How to Check If Your Site Is Included

Brass-SEO frames this as a check you can run yourself. Common Crawl publishes a searchable index of every URL it has captured, so you can look up your own domain and see which of your pages sit in the corpus that trains AI.

Open the Common Crawl URL index, search your domain, and read back the list of captured pages. A long list means your content has been feeding the archive for years. An empty result usually means one of two things: a young site the crawler hasn't reached, or a robots.txt that turned CCBot away. Either way, you now know where you stand instead of guessing.

How to Control CCBot

Brass-SEO treats CCBot access as a decision you own, not a default you accept. Common Crawl's crawler identifies as CCBot, and a site opts out of future crawls with a single robots.txt block that disallows that user agent, though snapshots already collected stay in the archive.

User-agent: CCBot
Disallow: /

Two limits are worth stating plainly. This stops future collection, not the copies already taken. And CCBot is only Common Crawl. OpenAI's GPTBot, Anthropic's ClaudeBot, and Google's crawlers are separate, each with its own robots.txt token. Blocking one leaves the others running. The Brass-SEO research index maps the full set of AI crawler controls so you can decide crawler by crawler rather than reaching for a blanket rule.

Should You Block It?

Brass-SEO's default recommendation is to stay in unless you have a specific reason to leave. For a business that wants customers to find it through AI, being in the corpus that models learn from is reach, and blocking CCBot trades a small, uncertain privacy gain for a real loss of visibility.

There are fair reasons to block. Proprietary research, paywalled material, a membership library you sell access to. For a marketing site whose whole job is to be found, hiding from the training data works against you. A cleaner move for most sites is to block training crawlers while allowing the search crawlers that power AI answers, so your content stays quotable even when it stays out of the next model. Brass-SEO's take on SEO versus GEO walks through that split in more detail.

Frequently Asked Questions

Is my website automatically in Common Crawl?

Most likely yes, if your pages are public and you have not blocked CCBot in robots.txt. Common Crawl has archived the open web since 2007 and adds 3 to 5 billion pages a month. You can confirm by searching your domain in Common Crawl's public URL index.

Does blocking CCBot remove my site from ChatGPT?

No. Blocking CCBot stops your pages from entering future Common Crawl snapshots, but it does not remove content already collected, and it does not affect other crawlers like OpenAI's GPTBot or Anthropic's ClaudeBot, which are separate. Controlling AI access means addressing each crawler on its own.

If I block AI crawlers, will I still show up in AI search answers?

Not necessarily. Many providers run separate crawlers for training and for live search. Blocking a training crawler such as CCBot or GPTBot while allowing the search crawler keeps your content out of the next model while staying eligible for AI answer citations. Brass-SEO's research index covers the split at /research/ai-crawler-control.

How much of GPT-3 came from Common Crawl?

Filtered Common Crawl supplied 60% of GPT-3's training mix and 410 billion tokens, the largest single source, according to the GPT-3 paper (Brown et al., 2020). The raw crawl was 45 terabytes of compressed text, reduced to 570 gigabytes after filtering.


Find Out If AI Can See Your Site

Being in the data that trains AI is the start. Being quoted by it is the goal. Brass-SEO reads your Google Search Console and Google Analytics data, both required, and scores your pages for AI Citability so you know which ones an AI engine can actually use. The plan is $25 a month with a three-day free trial, and setup takes about two minutes. Start a three-day free trial and see where your best pages stand.

Sources: Common Crawl · Brown et al., 2020 — Language Models are Few-Shot Learners (GPT-3) · Brass-SEO Research Index — AI Crawler Control

Ready to try Brass-SEO?

Get AI-powered SEO insights from your Google Search Console and Analytics data.