← Research

ChatGPT beats both Google AI Mode and AI Overviews in quality 84% of the time

A pilot study: blind, ranked comparison of ChatGPT, Google AI Mode and AI Overview on 25 real prompts. ChatGPT won 21 of 25. Full per-query data below.

By 5 min readAnswer Engine Optimization

What we did: took 25 real prompts, ran each one through ChatGPT, Google AI Mode and Google AI Overview, stripped the labels off, and had an independent judge rank the three answers.

  • ChatGPT came first on 21 of 25.
  • The two Google surfaces swap second and third from prompt to prompt — the gap between them is not stable.

#Summary

EngineMean overall rankFirsts
ChatGPT1.2021
AI Overview2.281
AI Mode2.523

#Method

#Takeaways

##ChatGPT wins by committing

  • It picks a brand, quotes a number, gives a name you can actually look up, says what the deciding factor is — and then stops.
  • The Google surfaces more often do competent analysis and then decline to land it, offering to keep going instead of answering. That cost them five prompts outright.
  • Being specific and occasionally wrong beat being thorough and non-committal.
  • That is what most people are asking for: 22 of the 25 questions (88%) wanted a specific answer — a name, a price, a pick, a yes or no. Only three were open-ended enough that a survey would do.

##More sources did not mean a better answer

  • Winning answers cited 5.0 sources on average. Second place cited 8.2, third 8.5.
  • The winner had the fewest sources of the three on 10 of 25 prompts, and the most on only 3.
  • On one prompt AI Mode cited 23 publishers — including a library-discovery platform and a job board — and still finished last.
  • Volume of sources is not a quality signal.

##The best answers are often not the longest

  • Winning answers averaged 4,259 characters, against 4,914 for second and 5,121 for third.
  • The winner was the shortest of the three on 13 of 25 prompts.
  • On one prompt the winning answer used 2,317 characters against the loser’s 5,982.
  • Length barely moves the ranking in either direction. Read this as length does not buy quality — not as short answers winning on their own.
  • Longer reads as more thorough and scores as more padded.

##If the answer is a name, leaving it out is fatal

  • Asked for Telegram channels, only ChatGPT gave handles that actually resolve. The Google answers named channels the user cannot find.
  • Same pattern as answers that describe an affiliate programme but withhold the commission rate.
  • The useful unit is a name someone can look up, not a description of it.

##Answering a prompt you should have questioned is worse than declining

  • One prompt was missing the context needed to answer it. AI Overview invented a scenario and shipped it with 7 citations; the other two asked what was meant.
  • Citations attached to a made-up premise make the error harder to spot, not easier.

##Most mistakes are staleness, not invention

  • Across the 25 prompts the errors were overwhelmingly out of date or mischaracterised — a real company described wrongly, a real product with last year’s terms, a rebrand not yet absorbed.
  • That changes the fix: publish fresher public information, rather than police a model for making things up.
  • One exception: a finance prompt produced four fund-return figures nobody can reproduce (26.46% claimed against 23.79% published) — in the one vertical where a wrong number costs money directly.

##Rebrands take months to land on Google

  • Only ChatGPT knew Velocity Global had become Pebl.
  • All three still said “LambdaTest” 8 months after the TestMu AI rebrand.
  • Budget many months for a rename to reach these surfaces.

##What all three missed is the opening

For every prompt the judge also recorded what none of the engines said. If no answer engine surfaces a fact, no competitor is being cited for it either. Three examples:

  • On a 7-year investment hold, tax and plan fees matter more than the difference between the funds being compared. None of the three raised either.
  • Suno’s free tier stopped allowing downloads a week after we captured. All three recommended it; none mentioned the cap.
  • Amazon runs its own vetted directory of service providers — the obvious place to check an agency before hiring one. No engine pointed to it.

#What this does and doesn’t prove

  • It’s a pilot, not a benchmark. 25 prompts, one geography (a US locale, with the Google answers pinned to New York). Trust the direction; treat the exact percentages as provisional.
  • Signed-out and free tiers only. None of this speaks to what a paying or logged-in user sees. ChatGPT Plus or Google AI Pro could rank differently.
  • A second judge agreed. Eight prompts were re-judged blind by a different model — one of Google’s own — which matched the first judge on seven of eight, and still put Google’s surfaces second and third every time. So the result is not the judge favouring its own family.
  • Refusing to answer counts against a surface. Several Google losses were for declining to conclude after solid analysis. If you think a search product should stay neutral, you would score those differently.
  • Judges can disagree on facts, not just taste. On one prompt two judges searched live and reached opposite conclusions about what was true.
  • The remaining 17 prompts were scored by one judge.

#Appendix: raw data

Every query, its full 1-2-3 order, and the three answers exactly as captured — citation counts and cited publishers included. Click a row.

ChatGPT
1.20
mean rank · 21 first places
AI Overview
2.28
mean rank · 1 first place
AI Mode
2.52
mean rank · 3 first places

Click a row for why the winner won, then Show raw model output, then any engine to read its answer exactly as captured.

ExpandQueryWinnerWhy it won

##Answer length and citations

#ChatGPT charscitesAI Mode charscitesAI Overview charscites
q15,974154,98435,88111
q27,9191011,9581610,90413
q33,38155,646116,71417
q44,26374,94343,7528
q52,85067,620254,5306
q62,31755,26185,98212
q73,92544,30635,67610
q83,19453,16123,4808
q92,71697,789175,17111
q104,18321,79513,89921
q113,03245,01337,26312
q12142022803,4337
q134,24147,865234,1179
q1410,90712,79302,0530
q153,11059,37895,95512
q164,23547,10954,88319
q171,98454,995143,4526
q182,56203,87101,6890
q196,16977,07022,9567
q203,22354,37744,89210
q216,12867,25767,47417
q222,25654,39012,3033
q234,76676,17455,34814
q243,04801,99504,5616
q254,87825,949143,6806
mean4,0564.95,4377.04,8029.8