AEO Presence Report

Is your website in ChatGPT's training data?

Enter a URL. ACME.BOT scans CommonCrawl and FineWeb, the public datasets behind today's largest LLMs, and returns a per-URL report.

Scans CommonCrawl (2B+ URLs) and FineWeb (200M+ URLs) · Free with an ACME.BOT account, no credit card.

Sample report

What you get back

Every URL from your domain found in CommonCrawl and FineWeb, organized by dataset with crawl date and page depth.
Sample report: URLs from a domain found in CommonCrawl and FineWeb, with crawl date and page depth.
  • 1

    Per-URL breakdown

    A verdict for every matched URL, not just the domain. You see exactly which pages made it in.

  • 2

    Dataset coverage

    The FineWeb and CommonCrawl columns show which dataset each URL appears in. FineWeb is the stronger training signal.

  • 3

    Freshness timestamps

    The month each URL was last crawled. Useful for spotting content that has dropped out of the index.

Why it matters

Why training-data presence matters in 2026

What ChatGPT, Perplexity, and Gemini say about your brand starts with whether your pages are in their training data. Without that, models guess, and they often get it wrong.

Upstream cause of AI answers

AI assistants form their picture of your business from training data first, then layer real-time search on top. What they learned up front still colors the answer.

Citations follow presence

Pages that sit in CommonCrawl and FineWeb turn up disproportionately often among what Perplexity and ChatGPT Search cite. Absence is the first gap to close.

Strategic diagnostic

Know where you stand and how much AEO content you need. Getting presence is a minimum. Getting citations is the goal.
How it works

How does the scan work?

ACME.BOT searches FineWeb and CommonCrawl records, notes where and when each URL was crawled, and lists every match.
1

You submit a domain

A root URL or a full path.

2

You sign in to ACME.BOT

The scan runs inside your account. Reports are free, no credit card.

3

ACME.BOT queries the indexed datasets

We search CommonCrawl and FineWeb. CommonCrawl holds close to 2 billion URLs and refreshes monthly. FineWeb is a subset of it, roughly 200 million URLs that survived LLM-grade filtering.

4

You get a per-URL report

Dataset, date last seen, and the matching URL for each hit. Export to CSV, or open the same data in ACME.BOT's Page Health View to track it over time.

Presence vs citations

Training-data presence vs. AI citations

Presence is what the model learned; citations are what it quotes today. Presence feeds citations, then freshness and structure decide which passages get pulled.
What it is
Presence (this report)
URLs found in training datasets
Citations (inside ACME.BOT)
Passages quoted by ChatGPT, Perplexity, and Gemini today
What it tells you
Presence (this report)
What the models learned
Citations (inside ACME.BOT)
What the models actually repeat
How to improve it
Presence (this report)
Get crawled: robots.txt, sitemap, fresh content, backlinks
Citations (inside ACME.BOT)
Publish AEO-structured content; refresh when citations drop

ACME.BOT tracks both: presence here, citations inside the product.

What to do next

What to do after you see the report

Can't find URLs in training data? Allow AI crawlers, publish new content regularly, and structure pages for easy AI extraction.

If most URLs are missing

Submit an updated sitemap, audit AI crawler access in robots.txt (GPTBot, CCBot, PerplexityBot), and get inbound links, since CommonCrawl prioritizes link-discovered pages.

If older URLs are present but new ones aren't

CommonCrawl runs monthly, but it still has to find the page through links. Keep refreshing and promoting new content.

If URLs are present but AI answers are still wrong

Presence alone is not enough. Freshness and structure decide what gets quoted, which is why AEO-optimized content matters.

If you want to track this over time

Each URL is assigned a Page Health View showing its search ranking, presence, and AI mentions by article.

Dataset coverage

Which datasets does ACME.BOT scan?

Two public corpora: CommonCrawl (used by GPT-3, GPT-4, Llama, and most major LLMs) and FineWeb (the filtered subset used for modern frontier training).
2B+
CommonCrawl

The largest open web crawl, refreshed monthly. A major training source for GPT-3, GPT-4, Claude, and Llama.

200M+
FineWeb

HuggingFace's quality-filtered subset of CommonCrawl, built with the heuristics frontier labs use. A good proxy for what today's training runs see.

No guarantees, but the strongest public signal there is

Presence in CommonCrawl or FineWeb doesn't guarantee inclusion in any specific LLM's training set, since most frontier labs mix these public datasets with unpublished proprietary data. But CommonCrawl is the backbone of nearly every open-corpus training run, and FineWeb mirrors the filters frontier labs apply internally. Absence from both is almost always a red flag; presence across both is the strongest public evidence that a model had access to your content.
FAQ

Frequently asked questions

Is my website in ChatGPT's training data?
What's the difference between CommonCrawl and FineWeb?
How often are these datasets updated?
Can I remove my content from existing training datasets?
How do I get my site included in future crawls?
Why are some of my pages present but not others?
Does training-data presence guarantee AI citations?
How is this different from Profound or Indexly?
Is this free?
How accurate is the scan?