Is your website in ChatGPT's training data?
Scans CommonCrawl (2B+ URLs) and FineWeb (200M+ URLs) · Free with an ACME.BOT account, no credit card.
What you get back

- 1
Per-URL breakdown
A verdict for every matched URL, not just the domain. You see exactly which pages made it in.
- 2
Dataset coverage
The FineWeb and CommonCrawl columns show which dataset each URL appears in. FineWeb is the stronger training signal.
- 3
Freshness timestamps
The month each URL was last crawled. Useful for spotting content that has dropped out of the index.
Why training-data presence matters in 2026
Upstream cause of AI answers
Citations follow presence
Strategic diagnostic
How does the scan work?
You submit a domain
A root URL or a full path.
You sign in to ACME.BOT
The scan runs inside your account. Reports are free, no credit card.
ACME.BOT queries the indexed datasets
We search CommonCrawl and FineWeb. CommonCrawl holds close to 2 billion URLs and refreshes monthly. FineWeb is a subset of it, roughly 200 million URLs that survived LLM-grade filtering.
You get a per-URL report
Dataset, date last seen, and the matching URL for each hit. Export to CSV, or open the same data in ACME.BOT's Page Health View to track it over time.
Training-data presence vs. AI citations
| Presence (this report) | Citations (inside ACME.BOT) | |
|---|---|---|
| What it is | URLs found in training datasets | Passages quoted by ChatGPT, Perplexity, and Gemini today |
| What it tells you | What the models learned | What the models actually repeat |
| How to improve it | Get crawled: robots.txt, sitemap, fresh content, backlinks | Publish AEO-structured content; refresh when citations drop |
ACME.BOT tracks both: presence here, citations inside the product.
What to do after you see the report
If most URLs are missing
Submit an updated sitemap, audit AI crawler access in robots.txt (GPTBot, CCBot, PerplexityBot), and get inbound links, since CommonCrawl prioritizes link-discovered pages.
If older URLs are present but new ones aren't
CommonCrawl runs monthly, but it still has to find the page through links. Keep refreshing and promoting new content.
If URLs are present but AI answers are still wrong
Presence alone is not enough. Freshness and structure decide what gets quoted, which is why AEO-optimized content matters.
If you want to track this over time
Each URL is assigned a Page Health View showing its search ranking, presence, and AI mentions by article.
Which datasets does ACME.BOT scan?
The largest open web crawl, refreshed monthly. A major training source for GPT-3, GPT-4, Claude, and Llama.
HuggingFace's quality-filtered subset of CommonCrawl, built with the heuristics frontier labs use. A good proxy for what today's training runs see.