research

We asked ChatGPT the same question 100 times. The top three never changed. Everyone else was a coin flip.

1,000 identical buying questions to ChatGPT, Gemini and Claude. The same three brands led every answer; below them a brand's odds swung from 1% to 62%.

Quick answer: we asked ChatGPT, Gemini and Claude the same three buying questions 100 times each — "best CRM for a small business", "best project management tool for a small team", "which accounting software should a freelancer use". 1,000 answers in one sitting, no personalization, one phrasing per question. The top of every answer was locked: HubSpot, Zoho and Pipedrive were recommended in 100 of 100 ChatGPT answers about CRMs, and HubSpot was named first in 100 of 100 — on all three engines. Below the top three, the answer is a coin flip: Close appeared in 13 of 100 and monday in 12, a brand's odds ran from 1% to 62% depending on the engine, and the locked three carried nearly all of the overlap between one answer and the next. ChatGPT read one affiliate listicle in 97 of its 100 CRM answers, and for the accounting question it did not search the web at all. The single "AI visibility score" this industry sells describes the locked top, where nobody needs it, and is noise in the tail, where almost every business lives. Every number below comes from the runs, and the raw answers are published.

Why ask the same question 100 times?

Because that is the only way to know what an AI recommendation is. Large language models are not deterministic in production. Thinking Machines Lab sampled one prompt 1,000 times at temperature zero — the setting that is supposed to make the model behave identically every run — and got 80 different completions. Ouyang, Zhang, Harman and Wang found that with ChatGPT's default settings, between 48% and 76% of coding tasks produced no two identical outputs across requests. Those studies were about facts and code. Nobody had pointed the same test at the thing an entire industry now sells a number for: which brands an assistant recommends when a buyer asks.

The industry has noticed the problem from the other side. In May, Digiday quoted Paul Dyer, CEO of the agency /prompt: "If you use three different tools and give them the same prompts, you get three different answers." The usual explanation is that the tools measure differently. We wanted to know how much of the disagreement is simply the engine disagreeing with itself.

What we did

Three questions a real buyer types, chosen from categories with many well-known brands and no relationship to us:

  1. "What's the best CRM for a small business?"
  2. "What's the best project management tool for a small team?"
  3. "Which accounting software should a freelancer use?"

Each question went 100 times to each of three engines on 22 September 2026 (UTC), from one location, through the APIs, with no conversation history and no account personalization:

  • ChatGPTgpt-4.1 through the Responses API with the web-search tool available, so the model could search if it chose to. It searched in 100 of 100 CRM answers, 100 of 100 project-management answers, and 0 of 100 accounting answers.
  • Geminigemini-flash-latest with Google Search grounding offered on every call. It used it in 1 of 300 answers.
  • Claudeclaude-sonnet-5 through OpenRouter, no web search.

As a control, we asked ChatGPT the CRM question another 100 times with search switched off, to separate the churn that comes from retrieval from the churn that comes from sampling.

A brand counts as recommended in an answer when it appears in a heading, a numbered or bulleted item, a table row or a bold lead — the places where an answer names its picks. A brand mentioned only in passing ("integrates with TurboTax") is not counted. Named-first is the brand at the top of the answer's list. Every extraction rule and the full answer set are in the methodology at the end.

The top of every answer is locked

The same three brands were recommended in 899 of the 900 answers across the three questions and three engines. The one miss: ClickUp dropped out of a single Gemini answer.

CRM for a small business ChatGPT Gemini Claude
HubSpot 100/100 100/100 100/100
Zoho CRM 100/100 100/100 100/100
Pipedrive 100/100 100/100 100/100
Salesforce 92/100 18/100 85/100
Freshsales 83/100 21/100 67/100
monday 12/100 97/100 17/100
Copper 0/100 62/100 17/100
Close 13/100 20/100 2/100
Streak 0/100 18/100 23/100
Less Annoying CRM 0/100 16/100 7/100
Project management for a small team ChatGPT Gemini Claude
Asana 100/100 100/100 100/100
Trello 100/100 100/100 100/100
ClickUp 100/100 99/100 100/100
Notion 77/100 100/100 100/100
Basecamp 75/100 64/100 11/100
monday.com 38/100 51/100 66/100
Linear 0/100 94/100 94/100
Jira 0/100 59/100 23/100
Zoho Projects 10/100 0/100 0/100
Accounting for a freelancer ChatGPT Gemini Claude
QuickBooks 100/100 100/100 100/100
FreshBooks 100/100 100/100 100/100
Wave 100/100 100/100 100/100
Zoho Books 99/100 86/100 56/100
Xero 93/100 49/100 100/100
Bonsai 20/100 99/100 73/100
HoneyBook 0/100 42/100 2/100
FreeAgent 6/100 0/100 0/100

Read the top rows and the bottom rows of each table as two different worlds. In the top rows, an assistant recommends the brand every time you ask; a tracking tool's "visibility score" for HubSpot would read 100% today, 100% tomorrow and 100% next month, and would tell HubSpot nothing. In the bottom rows, the same tool reports a number that moves by tens of points between two runs an hour apart, with nothing about the brand having changed.

Who is named first

"Recommended somewhere in the answer" is generous. The brand at the top of the list is the one most buyers try first, and here the lock is tighter still.

Question ChatGPT names first Gemini names first Claude names first
CRM HubSpot, 100 of 100 HubSpot, 100 of 100 HubSpot, 100 of 100
Project management ClickUp, 72 of 100 Asana, 91 of 100 Trello, 100 of 100
Accounting QuickBooks, 61 of 100 FreshBooks, 55 of 100 Wave, 85 of 100

Across 300 CRM answers on three engines, HubSpot was named first 300 times. Nothing about the sampling noise that produces 80 different Feynman biographies reaches the first line of a CRM recommendation. Where the first name does move, it moves between two incumbents: ClickUp or Trello, QuickBooks or Wave, FreshBooks or Wave.

Below the top three, it is a coin flip

Now the tail. Every brand outside the locked set is a probability, and for most of them it is a low one with a wide error bar. The 95% intervals below are what an honest tracking tool would have to print next to a "score" built from a handful of runs.

Brand, CRM question ChatGPT 95% interval
Freshsales 83% 74–89%
Close 13% 8–21%
monday 12% 7–20%
Brand, CRM question Gemini 95% interval
Copper 62% 52–71%
Salesforce 18% 12–27%
Freshsales 21% 14–30%
Streak 18% 12–27%
Less Annoying CRM 16% 10–24%

Two things to notice. First, a brand's odds depend on the engine far more than on the brand: Salesforce is recommended in 92% of ChatGPT's answers and 18% of Gemini's; Copper is in 62% of Gemini's and 0% of ChatGPT's. Second, the interval on a tail brand measured from ten runs — which is more than most tools use — is wider than the movement anyone would celebrate. A brand that goes "from 12% to 23%" between two weekly reports has, more often than not, gone nowhere.

Answer-to-answer overlap tells the same story in one number. Comparing the set of recommended brands in every pair of answers, ChatGPT's CRM answers overlapped by 86%, Gemini's by 68%, Claude's by 63%. The locked top three carry all of that overlap. Strip them out and the remainder of a typical answer is different from the remainder of the next one.

What ChatGPT read to decide

ChatGPT searched the web for all 100 CRM answers and all 100 project-management answers. What it read is the most useful table in this study, because it is the part a business can act on.

Source, CRM question Read in
croclub.com 97 of 100 answers
techradar.com 89
cloudspress.com 77
live-agent.cn 63
bizstackhub.com 56
interobservers.com 43
freshworks.com 42
beltstack.com 34
softabase.com 32
softwareconnect.com 24

Thirty-four distinct domains across the 100 answers, a median of seven per answer, and one of them in nearly every answer. croclub.com is The CRO Club, an affiliate review site published by Black & White Zebra; the page ChatGPT read is its list of enterprise CRMs, opened 97 times to answer a question about small businesses. Its own first pick is monday, which ChatGPT then recommended in 12 of 100 answers. Reading a page and following it are different things: the model's top three came from somewhere deeper than the page it fetched.

The project-management question read a different set — starters-tools.com in 96 of 100 answers, spotsaas.com in 90, techdealforge.com in 77, an Asana community forum thread in 73 — and again a small circle of listicle sites decided who was in the room. No vendor's own site made the top five for either question. This is the same pattern we found when we audited our own category: assistants assemble recommendations from third-party pages, and being on those pages is what moves the tail.

Search or no search: where the churn comes from

The control run answers a question the industry argues about. With web search off, ChatGPT's CRM answers still locked the same top three — HubSpot, Zoho and Pipedrive in 100 of 100 — and named HubSpot first every time. But the tail got wider: 17 distinct brands recommended against 7 with search on, answer-to-answer overlap of 76% against 86%, and Insightly, Less Annoying CRM and Keap appearing from memory where the searched answers never mentioned them.

So retrieval narrows the tail rather than widening it. When the model reads the same seven pages every time, it draws its supporting cast from those pages. When it answers from memory, it draws from everything it has ever read, and the sample is noisier. The locked top three is not a retrieval effect at all: it survives with the search tool switched off, which means it lives in the model, not in the pages.

One result surprised us. Given the search tool on every call, ChatGPT used it for 100 of 100 CRM answers, 100 of 100 project-management answers and 0 of 100 accounting answers. Same tool, same day, same phrasing style. The model decides per question whether it needs the web, and for "which accounting software should a freelancer use" it decided it did not.

This matters for measurement. A tool that counts citations — did the assistant link to your site — would have reported zero citations for every accounting vendor, QuickBooks included, and a naive reading of that is "nobody is cited". The honest reading is that nothing was fetched, so citation was not on the table. We wrote about this distinction in what grounding means; this is what it looks like in practice.

What this means if you run a business

If you are in the locked top three of your category, an AI visibility score tells you nothing. It will read 100% every week. Your risk is not the score moving; it is the small circle of listicle pages the engine reads changing their minds, and you should be watching those pages, not a dashboard.

If you are in the tail, one number is noise. The interval on a 12% brand from ten runs is roughly 2% to 40%. Any product that reports "you moved from 12% to 23%" without an interval is reporting sampling error as progress. Ask for the number of runs and the interval, and if the vendor cannot give them, the number is decoration. This is why SeeGeo reports rates over repeated runs and never sells an "AI rank".

The way into an answer runs through the pages it reads. Ten domains decided the CRM answer. For most categories the list is that short. Find it for your category — the assistants will tell you, if you ask them and keep the citations — and get onto it. That is the whole of off-site AI visibility work, and it is not mysterious.

Measure by sampling, not by asking once. Twenty repeats of a fixed question set, reported as a rate with an interval, is the minimum that turns an anecdote into a measurement. Our own free audit measures whether a site is readable and liftable — readiness — and our tracking asks the engines repeatedly for exactly this reason.

Frequently asked questions

Does ChatGPT give the same answer every time? For the first names in a recommendation, almost exactly: HubSpot was named first in 100 of 100 CRM answers on three engines. For everything below the top three, no: a tail brand appeared in anywhere from 1% to 62% of answers depending on the engine, and consecutive answers to the same question overlapped by only about 86% on ChatGPT.

Why do AI visibility tools disagree with each other? Partly because they measure differently, and partly because the engine disagrees with itself. Two tools sampling the same tail brand ten times each will get different numbers by chance alone; the 95% interval on a 20% brand from ten runs spans roughly 6% to 51%. Tools that report a single figure without an interval cannot be reconciled, because neither figure was ever precise.

Is this because ChatGPT searched the web? No. With search switched off, the same top three were still recommended in every answer and HubSpot was still first every time. Search made the tail narrower, because the model kept reading the same seven pages. The lock is in the model; the churn is in what surrounds it.

Which sites decide the answer? For the CRM question, croclub.com (97 of 100 answers), techradar.com (89) and cloudspress.com (77). For project management, starters-tools.com (96), spotsaas.com (90) and techdealforge.com (77). No vendor's own site appeared in the top five for either question.

Can I see the raw answers? Yes. All 1,000 answers, the extraction rules and the summary statistics are published at github.com/Tunisian-Aaron/ai-visibility-study-2026 under CC BY 4.0, alongside our 107-site audit study.

Full methodology

Published verbatim; runs conducted 22 September 2026 (UTC), from a single location, through the engines' APIs.

Questions. Three, chosen for well-known brand sets, buyer intent, and no relationship to SeeGeo: "What's the best CRM for a small business?", "What's the best project management tool for a small team?", "Which accounting software should a freelancer use?". One phrasing each; paraphrase variance is not measured here.

Engines and settings. ChatGPT: OpenAI Responses API, model gpt-4.1, web_search_preview tool available on every call, provider default temperature, no system prompt, no conversation history; whether the model searched is read from the response's web_search_call items. Gemini: gemini-flash-latest via the Generative Language API with google_search grounding offered on every call; an answer counts as grounded when grounding metadata is returned (1 of 300). Claude: anthropic/claude-sonnet-5 via OpenRouter, max_tokens 1024, no tools. Control: ChatGPT as above with the search tool omitted, CRM question only, 100 runs. Three concurrent requests per engine; each call retried up to three times on transport error; no call failed after retry.

Extraction. A dictionary of brand names and aliases per question (37 CRM, 45 project-management, 52 accounting entries), matched case-insensitively at word boundaries. Recommended: the brand appears in a heading, numbered or bulleted item, table row, or bold-led line of the answer. Named first: the first such brand in reading order. Passing mentions in body sentences are excluded from both. Every heading in every answer that matched no dictionary entry was reviewed by hand; genuine brands found that way (HoneyBook, Bonsai) were added before the final pass. Overlap: mean Jaccard similarity of the recommended-brand sets over all pairs of answers in a cell. Interval: Wilson score interval, 95%.

Sources. For grounded ChatGPT answers, the domains of the URLs in the response's citation annotations, de-duplicated per answer, www. stripped.

Limitations. One day, one location, one phrasing per question, English only, three engines and one model each; results describe these runs and nothing else. Model behaviour changes with versions and load, and the same experiment a month later may lock a different three. Gemini's grounding was offered but almost never used, so its answers are effectively from memory; Claude had no search at all; the engines are therefore not compared on equal footing and we do not rank them. Brand extraction is dictionary-based and can miss a brand named in an unanticipated form; the manual review of unmatched headings is the mitigation, not a guarantee. We did not measure answer quality, only which brands were recommended.

Want to know where you stand? SeeGeo audits your site for search and AI visibility, then tells you what to fix — in plain language.

Run a free audit