# SeeGeo

> SeeGeo is a free AI-visibility audit and tracking tool for small businesses. It checks whether AI systems — ChatGPT, Claude, Gemini, Google AI Overviews — can crawl, read, and cite a website, grades the site A–F, and explains what to fix in plain language. Audits take about 30 seconds, no signup. There is also a done-for-you service where SeeGeo's team implements the fixes.

This is the full-text companion to https://see-geo.com/llms.txt: every article and reference page on see-geo.com in one markdown file, generated at build time (2026-09-09) from the same sources as the HTML pages. Each page begins with a level-one heading and a Source line carrying its canonical URL. Any page can also be fetched alone as markdown by appending .md to its URL or sending `Accept: text/markdown`. The study's aggregate data is served as JSON at https://see-geo.com/data/ai-visibility-study-2026.json (described at https://see-geo.com/data, CC BY 4.0). French: https://see-geo.com/fr/llms-full.txt · German: https://see-geo.com/de/llms-full.txt.

- https://see-geo.com/
- https://see-geo.com/llms.txt
- https://see-geo.com/sitemap.xml

---

# AI visibility statistics, 2026
Source: https://see-geo.com/ai-visibility-statistics · Updated 2026-09-08

SeeGeo audited 107 real small-business websites (restaurants, trades, clinics, shops, professional services) in United States, United Kingdom, Ireland, Canada, Australia, France, Germany in August 2026. Sample of 120; 8 unreachable, 3 moved and 2 that blocked the checker were excluded and never enter a denominator. Convenience sample: the figures describe these 107 sites on this date, nothing more.

## Headline figures

| Question | Answer | Count |
|---|---|---|
| What share of small-business sites explicitly block at least one major AI crawler in robots.txt? | 2.8% | 3 of 107 |
| How many block any AI crawler counting wildcard rules too? | 2.8% | 3 of 107 |
| How many have no JSON-LD structured data anywhere? | 27.1% | 29 of 107 |
| How many carry structured data on the homepage? | 70.1% | 75 of 107 |
| How many homepages plainly say what the business does? | 35.5% | 38 of 107 |
| How many never say what they do? | 64.5% | 69 of 107 |
| How many serve an llms.txt file? | 25.2% | 27 of 107 |
| How many lose their content without JavaScript? | 2.8% | 3 of 107 |
| How many sit behind Cloudflare? | 29.0% | 31 of 107 |
| How many graded C or worse on SeeGeo's audit? | 94.4% | 101 of 107 |

## Blocking per crawler

| Crawler | Explicitly blocked | Effectively blocked | Share of sites |
|---|---|---|---|
| [GPTBot](https://see-geo.com/bots/gptbot) | 3 of 107 | 3 of 107 | 2.8% |
| [OAI-SearchBot](https://see-geo.com/bots/oai-searchbot) | 0 of 107 | 0 of 107 | 0.0% |
| [ClaudeBot](https://see-geo.com/bots/claudebot) | 3 of 107 | 3 of 107 | 2.8% |
| [PerplexityBot](https://see-geo.com/bots/perplexitybot) | 1 of 107 | 1 of 107 | 0.9% |
| [Google-Extended](https://see-geo.com/bots/google-extended) | 2 of 107 | 2 of 107 | 1.9% |
| [CCBot](https://see-geo.com/bots/ccbot) | 2 of 107 | 2 of 107 | 1.9% |
| [Googlebot](https://see-geo.com/bots/googlebot) | 0 of 107 | 0 of 107 | 0.0% |
| [Bingbot](https://see-geo.com/bots/bingbot) | 0 of 107 | 0 of 107 | 0.0% |

Of the 31 Cloudflare-fronted sites, 1 also block in robots.txt; of the 76 others, 2 do.

## Grade distribution

| Grade | Sites | Share |
|---|---|---|
| A | 0 | 0.0% |
| B | 6 | 5.6% |
| C | 38 | 35.5% |
| D | 38 | 35.5% |
| F | 25 | 23.4% |

Grades are anchored to this distribution: an A is the top decile of real sites, so a B is already a good result.

## By platform

| Platform | Sites | Block any AI crawler | No structured data |
|---|---|---|---|
| wordpress | 38 | 0 | 4 |
| custom/unknown | 34 | 2 | 20 |

Only platforms with enough sites for a fair comparison are shown; smaller groups are in the JSON with null figures.

## Method and license

Each site was crawled once by SeeGeo's audit engine the way an AI crawler reads it — honest user agent, no JavaScript, homepage plus up to about eight pages, robots.txt, sitemap and an llms.txt probe — and scored deterministically. "Blocks bot X" means robots.txt names X and disallows the site root; wildcard-only rules are counted separately as "effective". Individual businesses are never named. Aggregate generated 2026-08-17.

- Data (JSON): https://see-geo.com/data/ai-visibility-study-2026.json
- Dataset page with field dictionary: https://see-geo.com/data
- Repository with methodology and findings: https://github.com/Tunisian-Aaron/ai-visibility-study-2026
- Full write-up: https://see-geo.com/blog/how-many-websites-block-ai-crawlers
- License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
- Cite as: SeeGeo, AI visibility study of 107 small-business websites, 2026.

---

# Agent-to-agent commerce: what actually happened in the first year (and what it means for your business)
Source: https://see-geo.com/blog/agent-to-agent-commerce-what-happened · Updated 2026-09-08

> In 12 months agentic commerce launched, boomed and partly retreated: ACP, UCP, Instant Checkout's shutdown, the trust gap — and the layer that survived.

**Quick answer: agent-to-agent commerce — where your customer's AI agent transacts with your business's systems — went from announcement to production to partial retreat in about twelve months. The checkout layer sprinted and stumbled: OpenAI launched Instant Checkout in ChatGPT in September 2025, expanded it to all US users in February 2026, then shut it down in March 2026 after roughly 30 Shopify merchants (of a million promised) had actually integrated. The discovery layer, meanwhile, quietly boomed: Shopify's own telemetry shows AI-referred traffic up ~8x and AI-driven orders up ~13x year-over-year in Q1 2026, and Adobe measured AI-sourced retail traffic growing 393% and converting 42% better than organic by March 2026. The lesson of year one is the whole strategy: checkout mechanics keep churning, but being findable and machine-readable — the discovery layer — is where the volume, the conversions, and the durable work all live.**

This is the honest chronicle of agentic commerce's first year, with every claim graded the way [we grade everything](https://see-geo.com/blog/ai-visibility-glossary): established, emerging, or hype. It updates [our August piece on agentic shopping](https://see-geo.com/blog/geo-agentic-shopping), written before the checkout layer had been tested at scale. This space changed materially three times in twelve months; we re-verify the protocol status lines before every republish.

## What is agent-to-agent commerce, precisely?

The full vision: your customer tells their AI agent what they need ("running shoes for trail use, under $150, delivered by Friday"), and the agent interprets the intent, queries merchant catalogs, compares options, and completes the purchase — talking to *your* systems (your product feed, your policies, your checkout endpoint) rather than browsing your website like a human. When the merchant side is also automated — structured feeds answering structured queries, policies expressed in machine-readable form, an agent-compatible checkout — you have machines transacting with machines. The human sets parameters; software does the shopping.

How much of that exists today? The discovery half is real and scaled. The transaction half is real but wobbly. The negotiation half — your agent haggling with a merchant's agent — is still mostly a conference-slide vision. Let's take them in order of what actually happened.

## The twelve-month timeline (established facts, dated)

**September 2025:** OpenAI and Stripe launch the Agentic Commerce Protocol (ACP) — an open standard, Apache 2.0 licensed — alongside Instant Checkout in ChatGPT: US users buying from Etsy merchants in-chat, with over a million Shopify merchants announced as "coming soon." Merchants remain the merchant of record; OpenAI takes a transaction fee (reported at 4%) while stating that product recommendations are based on relevance, not enrollment.

**November 2025 – January 2026:** the land rush. Perplexity launches Instant Buy with PayPal; Salesforce announces ACP support; holiday data lands — Salesforce ties AI and agents to 20% of retail sales and $262 billion during the 2025 holiday season, and Adobe records 805% year-over-year growth in AI-driven retail traffic on Black Friday, with AI-referred visitors completing purchases at a 38% higher rate. (Grade these vendor figures as directional: "AI-influenced" is a broad, unaudited definition — but the direction is corroborated across independent platforms.) In January, Microsoft launches Copilot Checkout, and Google unveils the Universal Commerce Protocol (UCP) at NRF with co-development from Shopify, Etsy, Wayfair, and Target, and endorsements from 20+ partners including Stripe, Visa, Mastercard, and American Express.

**February 2026:** OpenAI expands "Buy it in ChatGPT" to all US users, including the free tier. The card networks go live in earnest through the spring — Mastercard completes its first live agentic transactions in Asia-Pacific markets, Visa's Trusted Agent Protocol goes commercial after piloting with 100+ partners, and American Express ships an agentic developer kit with purchase protection for registered AI-agent purchases. Shopify flips on Agentic Storefronts by default, syndicating eligible merchants' catalogs to ChatGPT, Google AI Mode, Microsoft Copilot, and Perplexity simultaneously.

**March 2026 — the retreat:** OpenAI sunsets Instant Checkout. The stated reason: the initial version didn't offer the flexibility they wanted, so merchants now use their own checkout while OpenAI focuses on product discovery. The revealing numbers around that decision: Forrester's Emily Pfeiffer put actual Shopify merchant integrations at roughly 30 — against the million promised — and Walmart, which had made about 200,000 products available in ChatGPT, found conversion rates three times *lower* for in-chat purchases than for shoppers redirected to Walmart's own site. ACP itself survives — it changed jobs from checkout to discovery infrastructure — and the ecosystem's center of gravity shifted to "discover in the chat, transact on the merchant's site."

Meanwhile, one giant sat the whole dance out: Amazon joined neither ACP nor UCP, building its own walled garden instead — Rufus, its shopping assistant, reportedly serves 300 million users and drove an estimated $12 billion in incremental sales in 2025.

## So is agentic commerce failing? No — the funnel just sorted itself

Read year one carefully and a clean pattern emerges: **every layer of the agent funnel matured at a different speed, and the order is instructive.**

**Discovery: established, and compounding.** This is where the verified platform telemetry lives — Shopify's ~8x traffic and ~13x order multiples are counted sessions on real infrastructure, not surveys. Adobe's Q1 2026 data (393% traffic growth, 42% conversion advantage) is the same class. People are letting AI *find and shortlist* things at enormous and accelerating scale.

**Transaction: emerging, and messier than the press releases.** Consumers use AI to decide, then prefer to buy in familiar places — a Semrush survey found only 22% had ever bought a product *inside* an AI tool, while about half had purchased somewhere after using AI in their research. The trust numbers explain the Walmart conversion gap: depending on the survey, only 4% to ~30% of consumers trust AI at the payment-authorization step, versus ~62–63% happily using it for comparison and discovery. And when people do trust an agent to buy, Bain found they trust retailer-owned agents three times more than third-party ones. The rails are being built ahead of the riders — normal for infrastructure, fatal for anyone who confuses rails with demand.

**Negotiation (true agent-to-agent): speculation with scaffolding.** Merchant-side selling agents, machine-readable offers, agent identity verification — the pieces are appearing (agentic feed products, bot-management vendors repositioning to let *legitimate* shopping agents through while blocking scrapers), but live agent-negotiates-with-agent commerce is not a 2026 reality for any normal business. The forecasts here are enormous — McKinsey sketches $3–5 trillion in redirected global retail spend by 2030, Gartner projects $15 trillion in agent-intermediated B2B purchases by 2028 — and they are forecasts, not measurements. Respect the direction; don't spend against the number.

## How do agents pick which businesses to transact with? (The part that decides winners)

Strip away the protocol politics and this is the question that matters, because whichever checkout standard wins, the *selection* step comes first — and it works nothing like the ad auction merchants are used to.

Agents choose programmatically. When an agent evaluates candidates for "birthday gift, skincare, under $75," it queries structured catalogs and scores what it can verify: data completeness, price, live availability, shipping and return policies it can actually read, reviews, and merchant trust signals. Practitioner data points in one direction: merchants with near-complete structured attributes get surfaced dramatically more; a business with a beautiful website and a thin machine-readable feed gets skipped for one with a plain site and a clean feed. Policies matter in a new way too — an agent qualifying merchants against "must allow 30-day returns" needs your returns policy as structured markup, not as prose on a page it half-parses. And notably: you can't buy your way in — recommendations in the major surfaces are relevance-based, which means the quality of your machine-readable presence is the primary lever. For small businesses, that's the most level playing field commerce has offered in decades.

The readiness gap is the opportunity: early-2026 research found 40% of e-commerce businesses still standardizing their product data for agents and another 33% not started at all. Nearly three-quarters of your competitors haven't done the unglamorous work.

## What should a small or medium business actually do?

The March retreat is your strategy guide: OpenAI itself pulled back from checkout to focus on *discovery* — which tells you which layer is durable. In order:

1. **Win the discovery layer first — it's the same GEO work.** Before any agent transacts with you, an AI system has to find you, classify you, and trust you. Crawler access, a plainly stated identity, [extractable content](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt), presence in cited sources — the entire visibility stack is the top of the agent funnel. This pays today (the 42%-better-converting AI-referred humans) and pre-qualifies you for whatever agents do tomorrow.
2. **Make your commerce facts machine-readable.** Product schema with real prices and live availability, structured shipping and returns policies, complete attributes on your top products. If you're on Shopify, much of this now ships by default via Agentic Storefronts — verify it's on and your data is complete rather than assuming ([our platform study](https://see-geo.com/blog/website-platform-ai-visibility-defaults) covers what each platform switches on for you).
3. **Let your platform carry the protocol risk.** ACP vs. UCP vs. whatever's next is a war between giants; Shopify, BigCommerce, Wix, and the payment networks are integrating the winners on your behalf. Don't hand-build protocol integrations; do keep the toggle on and the feed clean.
4. **Skip the hype spend.** No paid "agent optimization" service, no protocol consultancy, no panic replatforming. The two-question test still governs: cost if it does nothing, evidence it does something.

The one-sentence version of year one: **the machines learned to shop before the humans agreed to let them pay — and the businesses winning both eras are the ones machines can read.** That reading layer is what [SeeGeo's free audit](https://see-geo.com/geo-audit) measures — crawler access, machine-readable identity, structured data, extractability — in about 20 seconds, free, no signup for your score.

---

## Frequently asked questions

**What is the difference between agentic commerce and agent-to-agent commerce?**
Agentic commerce is the umbrella: AI agents researching, comparing, and increasingly transacting on a shopper's behalf. Agent-to-agent commerce is the fuller vision where the merchant side is also automated — the buyer's agent interacting with a merchant's structured feeds, policies, and checkout endpoints rather than a human-oriented website. Discovery-stage agentic commerce is established at scale in 2026; true agent-to-agent negotiation remains early.

**Did ChatGPT shopping fail?**
The in-chat checkout version was retired in March 2026 — after low merchant uptake (~30 Shopify integrations) and weak in-chat conversion (Walmart measured 3x lower than site checkout) — but ChatGPT shopping *discovery* is bigger than ever, and the underlying Agentic Commerce Protocol survives as infrastructure. The accurate summary: checkout retreated, discovery won.

**What is the difference between ACP and UCP?**
ACP (Agentic Commerce Protocol) is the open standard from OpenAI and Stripe, launched September 2025, now oriented toward discovery and merchant-side checkout. UCP (Universal Commerce Protocol) is Google's coalition protocol, announced January 2026 with 20+ partners, aimed at Google's AI surfaces. They are not yet interoperable; most SMBs should let their commerce platform handle both rather than integrating directly.

**How do AI agents decide which store to buy from?**
Programmatically: structured data completeness, price, live availability, machine-readable shipping and return policies, reviews, and trust signals — evaluated against the shopper's stated constraints. Ad spend doesn't enter it; recommendations on the major surfaces are relevance-based. The practical consequence: catalog and policy data quality is the primary competitive lever, which favors well-organized small merchants over big-but-sloppy ones.

**Will consumers actually let agents buy things for them?**
Slowly. Trust collapses at the payment step — surveys through 2026 put willingness to let AI spend autonomously between 4% and ~30%, versus ~62–63% comfortable using AI for comparison and discovery — and consumers trust retailer-run agents about 3x more than third-party ones (Bain). Expect bounded, chore-like purchases first (reorders, bookings, gifts within a budget) while high-consideration buying stays human-approved.

**Is any of this relevant if I'm a small local business, not an e-commerce store?**
Yes — the discovery layer is identical. Agents and assistants recommending local services run on the same machine-legibility: clear identity, consistent listings, readable policies and hours, structured data ([our local guide](https://see-geo.com/blog/local-business-ai-recommendations-guide) walks through it). The transaction layer (bookings via agents) will reach services through reservation and scheduling platforms the same way checkout reached retail through Shopify — your job now is being findable and classifiable when the recommendation happens.

---

# We gave google.com an F. Here's exactly how we score websites — and why yours means something
Source: https://see-geo.com/blog/how-we-score-ai-visibility · Updated 2026-08-21

> Six categories, two pillars, grades anchored to 107 real small-business sites, and rules that stop errors reading as invisibility. The full scoring system.

**Quick answer: SeeGeo's audit crawls your site once, politely, the way an AI crawler does — no JavaScript execution — and scores it across six categories: crawler access, technical foundation, structured data, content extractability, off-site presence, and entity clarity. Categories combine into two pillar scores (classic SEO readiness and GEO — readiness to be cited in AI answers), which average into a 0–100 number and a letter grade. The letters are anchored to reality: an A means the top tenth of the 107 real small-business websites we measured in [our published study](https://see-geo.com/blog/how-many-websites-block-ai-crawlers), not an arbitrary cutoff. The score is deterministic — same site in, same score out — every report states exactly what was and wasn't evaluated, and a provider outage can never read as your invisibility.** The rest of this post is the whole rubric, including the parts that make grades go down, the parts that deliberately can't, and why an F for google.com is both correct and meaningless.

Scoring tools usually treat their formula as a trade secret. We think that's backwards: a score you can't interrogate is a score you can't trust, and [trust is the entire product](https://see-geo.com/methodology). So here's ours.

## The crawl: we read your site the way AI does

Every audit starts with one polite crawl — your robots.txt, your sitemap, your homepage, and a budgeted set of key pages, fetched with our honest user-agent and **without executing JavaScript**. That last choice is the load-bearing one: [Vercel and MERJ measured](https://vercel.com/blog/the-rise-of-the-ai-crawler) that the major AI crawlers download JavaScript but don't run it, so a page whose content only exists after scripts run is, to most AI systems, an empty room. We deliberately see what they see. If your content survives our crawl, it survives theirs.

## The six categories

**1. Crawlability & access — the gate.** Can crawlers get in at all? We evaluate your robots.txt against sixteen named crawlers individually — classic search bots and the AI roster (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, and friends) — plus bot-walls, noindex directives, and whether your security layer answers instead of your site. This category is special: a critical failure here doesn't just lower a number, it **caps your whole grade**, because nothing else matters if machines can't read you. The cap is proportional — blocking one minor AI crawler costs you far less than walling out everything — and if the blocking looks deliberate (the pattern big publishers use on purpose), the report says so instead of scolding you.

**2. Technical foundation.** Page speed, mobile viewport, HTTPS hygiene, clean URLs, sitemap health. Classic, unglamorous, still counted.

**3. Structured data.** The schema.org markup that lets a machine assert facts about you instead of inferring them — LocalBusiness, Organization, FAQ, Product. In our study, 27.1% of real small-business sites had none at all.

**4. Content extractability.** Could an AI *quote* you? Clear headings, real text (not text baked into images), answerable questions, statistics and sources. This is where the [Princeton GEO research](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt) says most citation variance lives, so it carries the most GEO weight.

**5. Off-site presence.** Wikipedia and mention analysis — does the wider web know you exist? Honesty note that matters: this category's data sources are still growing, so its weight in your score **scales with how many sources actually answered**. We'd rather a category count for less than pretend one lookup is a full picture.

**6. Entity clarity.** Can a machine state, with confidence, what you are, where you operate, and for whom? Our classifier tries — and if it can't classify you, neither can an LLM. This was the most-failed check in our study: 64.5% of sites never plainly say what the business does at the top of the homepage. It's also the cheapest thing on this list to fix.

You'll notice what's *not* a category: which AI engines you've paid to appear in (not a thing that exists), your follower counts, or a secret sauce coefficient. Volume of honest signal, nothing else.

## Two pillars, one grade — anchored to real sites

Each category feeds two pillar scores with different weights: **SEO** (search readiness — access and technical matter most) and **GEO** (AI-answer readiness — extractability and off-site matter most). The pillars average into your combined score, and the combined score becomes a letter.

Here's the part we changed recently, and why your letter is defensible in a way most tools' letters aren't: **the grade cuts come from the score distribution of 107 real small-business websites** — restaurants, trades, clinics, shops across seven countries — from [our published study](https://see-geo.com/blog/how-many-websites-block-ai-crawlers). An **A means the top decile** of real sites. B is the top quartile. C clears the median. Under our old arbitrary cutoffs, the best site we ever measured graded B and nothing could earn an A — a ruler nobody could calibrate against. Now the letter answers a real question: *compared to actual businesses competing for the same AI answers, where do you stand?*

## The rules that protect you from us

A scoring system reveals its character in its failure modes. Ours has four rules that all bend the same direction:

**An error can never read as invisibility.** If Wikipedia times out mid-audit, we retry; if it's truly unreachable, the report says "source unanswered" — it never says "no Wikipedia presence" because our network hiccuped. The same rule runs through our tracking product: failed API calls are flagged and excluded from every statistic.

**Unevaluated never counts against you.** If a category couldn't be measured, it's removed and the weights renormalize — and the report says so, right under the grade: *"Scored on 5 of 6 categories"*, with the reason. A screenshot of your grade carries its own caveat.

**Same site in, same score out.** The score is deterministic — no LLM judgment folded into the number, ever. When we do show you AI-generated readings (the "How an AI reads your site" panel, where a language model reads your homepage exactly as a crawler sees it), they're labeled, and they sit *beside* the score, never inside it. A tracked score only means something if the ruler holds still.

**The ruler is versioned.** When we improve the engine — new checks, re-anchored grades — the version number on your report changes and the change is documented. If your grade moves between audits, you can tell whether your *site* changed or our *measurement* did. We built this after watching a site grade differently two days apart and realizing we couldn't prove which had moved.

## So why does google.com get an F?

Because the audit measures one thing: **whether a machine that has never heard of a business can learn what it is from its website.** Google.com scores a perfect 100 on crawler access — and an F overall, because its homepage is a search box: no self-description, no structured data about the business, nothing to extract. Every one of those findings is *true*. And every one is *meaningless as a verdict on Google*, because household names live in every AI model's training data and get recommended no matter what their homepage says.

That's exactly the problem an unknown business doesn't have the luxury of. If you're not famous, the machines learn who you are from your website or they don't learn it at all — which is who this rubric is built for. (The report says all of this in context when you audit a famous site, before showing the grade. Skeptics testing us with google.com first: we see you, and fair enough.)

## What the score doesn't claim

The audit is a measurement of *readiness*, not a promise of *outcomes*. GEO is a [young discipline](https://see-geo.com/blog/ai-visibility-glossary); our recommendations reflect current research and practitioner consensus, not settled science, and every report says so verbatim. The audit also can't see whether AI engines actually mention you today — that's what [tracking](https://see-geo.com/pricing) measures, with repeated live queries across engines, because single answers are noise and trends are signal.

Run it yourself — [it's free, about 30 seconds, score with no signup](https://see-geo.com/#audit). And if you disagree with a finding, the evidence is printed under it: the robots.txt line, the missing markup, the page we fetched. Argue with the evidence, not with a black box. That's the point.

## Frequently asked questions

**How does SeeGeo calculate its website score?**
One JavaScript-free crawl feeds six category scores (crawler access, technical, structured data, extractability, off-site presence, entity clarity), which combine into SEO and GEO pillar scores and average into a 0–100 number. Letter grades are percentile-anchored: A = the top tenth of the 107 real small-business sites in our published study.

**Why did a famous website score badly on your audit?**
Because the audit measures whether a machine can learn what a business is from its website — the problem unknown businesses have and famous ones don't. A household name's homepage is often a pure application with no self-description, which grades honestly low while meaning nothing about the brand's actual AI visibility. Famous-site reports say this in context before the grade.

**Is the score affected by AI randomness?**
No. The audit score is fully deterministic — same site in, same score out — with no LLM judgment inside the number. AI-generated readings appear beside the score, clearly labeled. Our tracking product handles AI answer randomness separately, by measuring across repeated runs.

**What does "scored on 5 of 6 categories" mean?**
A category that couldn't be evaluated — say, an off-site data source didn't answer — is excluded and the remaining weights renormalize, so a missing data source never punishes your site. The report states which category was skipped and why.

**Can I improve my grade, and how fast?**
Usually, yes — the report is a prioritized fix list, impact-first, and the most common failures (no plain-language identity statement, no structured data) are cheap to fix. Re-audit any time for free; the engine version on each report tells you the comparison is apples-to-apples.

---

# When AI does the shopping: why GEO matters more in the age of agentic commerce
Source: https://see-geo.com/blog/geo-agentic-shopping · Updated 2026-08-19

> AI agents are starting to browse, compare, and buy for customers. What's real vs early in agentic commerce — and why GEO becomes your storefront.

**Quick answer: agentic shopping is the next phase of AI-driven commerce — AI agents that don't just recommend businesses but act for the customer: browsing sites, comparing options, filling carts, and completing purchases. The payment rails for this are being built right now (OpenAI's agentic checkout work with Stripe, Google's AP2 agent-payments protocol, card-network initiatives from Visa and Mastercard), though consumer adoption is still early. What it means for your business is simple and slightly unnerving: when the shopper is software, your AI visibility is your storefront. GEO — being findable, legible, and usable by machines — stops being a marketing channel and becomes the door the customer's agent either walks through or doesn't.**

*Update, September 2026: the checkout layer this post anticipated has since launched, scaled and partly retreated. [The first-year chronicle](https://see-geo.com/blog/agent-to-agent-commerce-what-happened) has the dated timeline and the protocol status; the advice below still holds.*

An honest tag before we start, in keeping with [how we grade everything](https://see-geo.com/blog/ai-visibility-glossary): agentic commerce is **Emerging**, not Established. The infrastructure announcements are real; the mass consumer behavior isn't here yet. This post is about why the preparation is worth doing *now* anyway — and why almost all of it is work you should be doing for today's AI recommendations regardless.

## What actually happens when an AI agent shops?

Today's AI shopping journey mostly ends at a recommendation: a customer asks ChatGPT for the best option, reads the answer, and clicks through to buy like a normal human. Agentic shopping extends the machine's role through the whole funnel. The customer says "find me a birthday cake for Saturday, under $60, near me, order it" — and the agent searches, shortlists, checks availability and price *by reading the actual websites*, and completes the transaction through an agent-compatible checkout.

Every step of that journey is a machine reading and operating your web presence:

- **Discovery** is AI retrieval — the same recommendation mechanics GEO already addresses ([query fan-out](https://see-geo.com/blog/query-fan-out-explained), citations, entity clarity).
- **Evaluation** is an agent parsing your pages: prices, availability, options, policies. If your price list is a JPEG and your hours live only in an Instagram bio, the agent's comparison table has blanks where your business should be — and agents don't squint at images the way patient humans do.
- **Transaction** is the genuinely new layer: protocols that let an agent pay on the customer's behalf, with the announced rails coming from OpenAI/Stripe, Google's AP2, and the card networks. This layer is where the "early" caveat lives — announced, piloted, not yet mass behavior.

One piece of this is already switched on without most merchants noticing: Shopify's Catalog syndicates eligible product data to agentic storefronts like ChatGPT by default — a second data path that robots.txt doesn't govern, [as we documented in our platform study](https://see-geo.com/blog/website-platform-ai-visibility-defaults). If you sell on Shopify, agents can already see your products; the only question is what they find when they look.

## Why does this raise the stakes for GEO specifically?

Because each step *removes a place where human forgiveness used to save you.*

A human shopper who half-remembers your bakery might type your name into Google despite your invisible website. An agent won't — if you're not retrievable, you're not on the shortlist. A human who lands on your confusing site might dig for the price. An agent scores what it can parse and moves on. A human might phone to ask if you deliver. An agent checks your structured data, finds nothing, and marks delivery: unknown — which, in a ranked comparison, is a no.

The recommendation era made AI visibility important; the agentic era makes it *load-bearing*. And the early conversion data suggests what's at stake: AI-referred shoppers already convert dramatically better than organic traffic — Adobe's 2026 panel found AI-assistant visitors converting 42% better than non-AI traffic, and Semrush's cross-industry analysis puts AI-driven visitors at roughly 4.4x organic conversion ([our full breakdown of that data](https://see-geo.com/blog/ai-visibility-revenue-data-2026)) — because the comparison happens before the click. Agentic shopping is that same pre-qualification taken to its endpoint: the agent arrives *ready to transact*. The businesses in its consideration set split the highest-intent demand that exists; everyone else splits nothing.

## What makes a website "agent-ready"? (The unglamorous truth)

Here's the part that should reassure you: agent-readiness is not a new exotic discipline. It's the same fundamentals GEO already demands, enforced more strictly:

**1. Access.** Agents fetch your pages with their own user-agents — the retrieval and user-triggered fetchers, the ones that act at the moment of a customer's request. Blocked crawlers and CDN bot-walls now don't just cost you a citation; they cost you a transaction. Worth knowing the distinction [our crawler library](https://see-geo.com/bots) covers: training bots (like GPTBot) shape long-term model knowledge, while retrieval and user-fetch bots (OAI-SearchBot, ChatGPT-User) power live answers and agent actions — blocking the second category is the expensive mistake.

**2. Content that survives without JavaScript.** The best-documented platform-level risk in AI visibility — [Vercel and MERJ's finding](https://vercel.com/blog/the-rise-of-the-ai-crawler) that major AI crawlers fetch JavaScript but don't execute it — applies doubly to agents doing live evaluation. If your prices render client-side only, an agent's read of your store may be an empty room.

**3. Machine-readable commerce facts.** Product schema with real prices and availability, services and policies as text, structured hours, LocalBusiness markup. This is the difference between an agent *knowing* your delivery cutoff and *guessing* it. In agentic comparison, complete structured data is what a good shelf position used to be.

**4. Semantic, operable HTML.** Real buttons, labeled forms, sane heading structure — the accessibility fundamentals. Agents navigating sites succeed on the same markup that screen readers do; agent-hostile design and accessibility debt are the same debt. (A pleasant side effect: fixing it serves human customers with disabilities today, whatever agents do tomorrow.)

**5. Entity clarity.** The agent's first job is classifying you — what you sell, where, for whom. The [plain-language identity statement](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt), consistent everywhere, remains the cheapest high-leverage fix in the entire stack — and the one [64.5% of small-business sites in our study](https://see-geo.com/blog/how-many-websites-block-ai-crawlers) get wrong.

Notice what's absent: no secret agentic meta-tags, no "agent optimization" services, no new file format to buy. If someone sells you agentic-commerce optimization as proprietary magic in 2026, apply the standard test — what's the cost if it does nothing, what's the evidence it does something — and keep your money.

## What should a small business actually do, and when?

**Now (because it pays today, agents or not):** the five fundamentals above. Every one of them improves your current AI recommendations and your human conversion while pre-building agent-readiness. This is the rare strategic bet with no downside branch — the work is identical whether agentic shopping arrives fast or slow.

**Watch (quarterly, not daily):** which checkout protocol wins, when the platforms you're on ship agent-checkout support beyond what's already live, and the first credible data on agent-completed purchase volume. When your platform offers an agent-compatible checkout toggle, turning it on should be boring — because everything upstream was already ready.

**Skip:** rebuilding your site "for agents," buying protocol-specific integrations before a standard wins, and any vendor promising placement in agent shortlists. The shortlist isn't for sale; it's computed from the fundamentals.

The honest one-line summary: agentic shopping is early, the rails are real, and the preparation is free — because it's the same GEO work that earns you AI recommendations today. The businesses that treated machine-legibility as infrastructure will wake up one quarter to find agents transacting with them; the ones that didn't will be reading think-pieces about it.

Want to know where you stand today? [SeeGeo's free audit](https://see-geo.com/#audit) checks the agent-relevant fundamentals — crawler access, no-JS content survival, structured data, entity clarity — in about 30 seconds, no signup for your score.

## Frequently asked questions

**What is agentic shopping?**
Shopping where an AI agent acts on the customer's behalf across the funnel — searching, comparing options by reading websites directly, and completing the purchase through agent-compatible payment rails. It extends today's AI recommendations (where the human still clicks and buys) into machine-executed transactions.

**Is agentic commerce actually happening in 2026, or is it hype?**
The infrastructure is real — OpenAI's agentic checkout work with Stripe, Google's AP2 payments protocol, and card-network initiatives are announced and piloting, and Shopify already syndicates product data to agentic storefronts by default. Mass consumer adoption is not here yet. The honest status is Emerging: worth preparing for with work that pays off today anyway, not worth panic-rebuilding for.

**How do I make my website ready for AI shopping agents?**
The same fundamentals that win AI recommendations, enforced strictly: allow the retrieval-class AI crawlers, ensure your content and prices are readable without JavaScript, publish complete structured data (Product, LocalBusiness, real prices and availability), use semantic accessible HTML, and state plainly what you sell and where. No exotic new formats are required.

**Will AI agents replace human shopping?**
Unlikely to replace it wholesale — the realistic near-term pattern is agents handling constrained, chore-like purchases (reorders, bookings, gift logistics within a budget) while humans keep the browsing they enjoy. But even partial adoption concentrates the affected demand onto the businesses agents can actually see and use.

**Does being agent-ready help with anything before agents arrive?**
Yes — that's the point. Every agent-readiness fundamental is also a today-fundamental: crawler access and extractable content drive current AI citations, structured data feeds rich results, semantic HTML is accessibility, and entity clarity converts human visitors. The agentic era doesn't ask for new work; it raises the price of skipping the old work.

---

# How to ask ChatGPT about your own business: the 15-minute self-check every owner should run quarterly
Source: https://see-geo.com/blog/check-what-chatgpt-says-about-your-business · Updated 2026-08-18

> The exact prompts to ask ChatGPT, Perplexity, and Gemini about your own business — and how to read the answers. A 15-minute quarterly check, step by step.

**Quick answer: you can find out exactly what AI assistants tell customers about your business in about 15 minutes, using four types of prompts asked across ChatGPT, Perplexity, and Gemini — and repeating each one a few times, because AI answers change between runs.** What you learn falls into four patterns (described wrong, competitor recommended instead, not known at all, or accurately recommended), and each pattern points to a different fix. This guide gives you the exact prompts, what to look for in the answers, and what each result means — no tools required.

Most owners have Googled themselves. Almost none have ChatGPT'd themselves — even though a growing share of their customers ask AI assistants for recommendations before they ever reach a search results page. This is the new "Google yourself," and it takes a quarter of an hour.

## Why check what AI says about you?

Because AI assistants are now answering questions about your business whether you participate or not. When someone asks "is [your business] any good?" or "best [your category] near me," the assistant composes an answer from whatever it knows and retrieves — your site, directories, reviews, forum threads, articles. If that picture is wrong, stale, or empty, the AI presents its version confidently to a customer you never see. The only way to know which version of you it's serving is to ask it yourself.

## Before you start: two rules that make the check valid

**Rule 1 — Ask each prompt 3 to 5 times, in fresh chats.** AI answers are [non-deterministic](https://see-geo.com/blog/ai-visibility-glossary#non-determinism): the same question produces different answers on different runs. Ask once and you have an anecdote; ask five times and you have a pattern. (You'll see this yourself by run three — the brand lists shuffle. That's normal, and it's why any tool selling you a fixed "AI rank" should be treated with suspicion.) Use a new conversation for each run so earlier answers don't contaminate later ones.

**Rule 2 — Ask like a customer, not like the owner.** Log out or use a browser profile that isn't soaked in your own brand searches where you can. Phrase prompts the way a stranger would — "best plumber in Rochester that does emergency calls," not your internal category jargon.

## The four prompts (run each on ChatGPT, Perplexity, and Gemini)

### Prompt 1: The identity check — "What is [your business name] in [city]?"

You're testing whether the AI *knows you exist* and describes you correctly. Read the answer for: right category? right location? current offerings, or the menu you retired two years ago? anything invented (a hallucinated founding date, a service you don't offer)?

### Prompt 2: The recommendation check — "What's the best [your category] in [your city/area]?"

The money question — the one your customers actually ask. Note every business named, in a simple tally across your runs. You're measuring three tiers: are you [*mentioned*](https://see-geo.com/blog/ai-visibility-glossary#mention) at all, [*cited*](https://see-geo.com/blog/ai-visibility-glossary#citation) (linked), or actively *recommended*? Also note who shows up instead of or alongside you — that's your real competitive set in this channel, and it often differs from who you think your competitors are.

### Prompt 3: The comparison check — "Alternatives to [the competitor who came up most]" and "[your business] vs [that competitor]"

This reveals how the AI positions you when a customer is actively choosing. Does it know your strengths? Does it repeat one old review's complaint as if it were your defining trait? Comparison prompts are where inaccurate or thin information hurts most, because the customer is at the decision point.

### Prompt 4: The trust check — "Is [your business name] trustworthy? What do reviews say?"

AI assistants synthesize review sentiment into a verdict. You're checking which sources it leans on (Google reviews? Yelp? a forum thread from 2021?) and whether the summary matches your actual current reputation. On Perplexity especially, note the cited sources — those exact pages are where your reputation lives, in the AI's eyes.

## How to read your results: the four patterns

**Pattern A — Described wrong or out of date.** The AI knows you but has stale or invented details. Fix: [entity](https://see-geo.com/blog/ai-visibility-glossary#entity) cleanup. Make your homepage state plainly what you are, update your About page, and align every directory and profile to one consistent name, category, and description. Inconsistency across the web is what blurs the AI's picture of you.

**Pattern B — Competitor recommended instead.** You exist but lose the recommendation. Fix: find out where the AI's answers come from — note the sources it cites for your category prompts, open them, and check whether you appear in those pages. Being absent from the listicles, directories, and review platforms the AI already cites is the most common and most fixable cause. Full article: [Why does ChatGPT recommend my competitor instead of me?](https://see-geo.com/blog/chatgpt-recommends-competitor-not-me)

**Pattern C — Not known at all.** The AI has effectively never heard of you. Check the plumbing first — can AI crawlers access your site? — then build the basics: a clearly-described website, presence on major directories and review platforms, and content that answers the questions your customers ask. Full article: [Is your website visible to ChatGPT?](https://see-geo.com/blog/is-your-website-visible-to-chatgpt)

**Pattern D — Accurately recommended.** Congratulations — now protect it. Note the sources being cited on your behalf and keep them healthy (fresh reviews, current listings), because the AI's picture of you is only as durable as its sources.

## Make it quarterly (and write it down)

A single check is a snapshot. The value compounds when you repeat it: keep a simple sheet — date, prompt, engine, mentioned/cited/recommended, how described — and run the same prompts every quarter. What you're building is a trend line, and the trend is the only honest measurement in a channel where individual answers vary. Fifteen minutes, four times a year, and you'll know something most of your competitors don't: what the machines that increasingly introduce businesses to customers actually say about yours.

If you'd rather not do it by hand: this exact loop — same prompts, multiple runs, across engines, tracked over time — is what [SeeGeo](https://see-geo.com/#audit) automates, alongside the site-side checks (crawler access, content structure) that the manual method can't see. But run the manual check at least once regardless. Reading the AI's own words about your business, live, is worth more than any dashboard for making the channel feel real.

## Frequently asked questions

**What should I ask ChatGPT to see what it knows about my business?**
Four prompts, asked like a customer: "What is [business name] in [city]?" (identity), "Best [category] in [city]?" (recommendation), "[Business] vs [competitor]" (comparison), and "Is [business] trustworthy — what do reviews say?" (trust). Run each 3–5 times in fresh chats on ChatGPT, Perplexity, and Gemini, and tally the pattern rather than trusting any single answer.

**Why does ChatGPT give different answers each time I ask?**
AI answers are non-deterministic by design — the same question legitimately produces different responses across runs. That's why a valid self-check repeats each prompt several times and reads the pattern, and why single-run "AI rankings" are noise.

**ChatGPT described my business incorrectly. Can I fix it?**
Yes, indirectly. AI systems learn from your website and everything published about you, so the fix is making the correct information unmissable: a plain-language description at the top of your homepage, an accurate About page, and consistent name/category/location across every directory, profile, and review platform. Corrections propagate over weeks to months depending on the engine.

**Can I pay to be recommended by ChatGPT?**
No — there's no buyable placement in the organic answers of the major assistants as of 2026. Recommendations come from being accessible to their crawlers, clearly described, and present in the sources they cite. That's earned, not bought — which is bad news for shortcuts and good news for small businesses that do the work.

**How often should I run this check?**
Quarterly as a baseline, plus after anything that changes your public footprint: a rebrand, a move, a site migration, a review surge, or a burst of press. AI models and their sources update continuously, so an annual check misses the drift.

---

# Your Google Business Profile is now an AI data source: the local business guide to being recommended by ChatGPT
Source: https://see-geo.com/blog/local-business-ai-recommendations-guide · Updated 2026-08-18

> 45% of consumers now ask AI for local recommendations (BrightLocal 2026). Your profile, reviews, and listings are AI data sources — six levers to win.

**Quick answer: AI assistants have become a mainstream way people find local businesses — [BrightLocal's Local Consumer Review Survey 2026](https://www.brightlocal.com/research/lcrs-ai-trust/) found AI use for local business recommendations jumped from 6% to 45% in a single year, making it a top-three discovery channel behind only Google and Facebook. The good news for local owners: the assets you already maintain for Google Maps — your Business Profile, reviews, directory listings, and website — are exactly what AI systems draw on. The work isn't new; the stakes are. This guide covers how AI answers "best [category] near me," where it pulls its information, and the six things that determine whether it says your name.**

## How do AI assistants actually answer "best pizza near me"?

Not from a single database — from a synthesis. When someone asks ChatGPT, Perplexity, or Gemini for a local recommendation, the assistant retrieves and combines several kinds of sources: local directory and profile data (your Google Business Profile and equivalents), review platforms and the *text* of reviews, "best of" articles and local listicles, community threads (Reddit is disproportionately cited), and your own website. Then it composes a short answer naming typically one to five businesses, often with a sentence of description each.

Two properties of that process matter enormously for a local owner. First, the assistant describes you in *its own words, assembled from others' words* — the language in your reviews and listings literally becomes the language of your recommendation. Second, answers vary between runs, so your goal isn't a fixed "rank" — it's appearing in as high a share of relevant answers as possible.

## Why this channel rewards local businesses specifically

Local is where AI recommendations are most decisive: the questions are high-intent ("emergency plumber near me open now"), the answer set is small, and the user typically acts on the recommendation rather than continuing to research. It's also where small operators can genuinely beat bigger ones — AI systems reward clear, consistent, well-reviewed [entities](https://see-geo.com/blog/ai-visibility-glossary#entity), not marketing budgets. A five-person business with immaculate listings, current information, and detailed reviews regularly out-appears a regional chain with a neglected footprint.

## The six things that determine whether AI says your name

### 1. One identity, everywhere (NAP consistency, promoted to survival trait)

Name, address, phone, category, hours — identical across your website, Google Business Profile, Apple Business Connect, Yelp, Facebook, and every directory that lists you. This was always local-SEO hygiene; for AI it's more fundamental, because the assistant is cross-referencing sources to decide whether "Joe's Plumbing," "Joe's Plumbing & Heating," and "Joes Plumbing LLC" are one business or three. A blurry entity doesn't get confidently recommended. Do the boring audit: list everywhere you appear, fix every discrepancy, kill duplicate listings.

### 2. A homepage that says what you are, where you are, in plain words

Within the first screen of your website: "[Name] is a [category] in [city/area] serving [whom] — [core services]." Not a slogan doing that job by implication. Your site is a primary source for what you *are*; if a machine can't classify you from the top of your homepage, you're relying on everyone else's descriptions of you — and hoping they're right. (In [our 107-site study](https://see-geo.com/blog/how-many-websites-block-ai-crawlers), 64.5% of small-business sites failed exactly this check — it's the most common gap there is, and the cheapest to fix.)

### 3. A complete, current Google Business Profile — treated as content, not a checkbox

Every field filled: precise categories (primary and secondary), services/menu with real descriptions, current hours (including holidays), photos, and the attributes that match how people actually ask ("wheelchair accessible," "open late," "free estimates"). Assistants answer qualified questions — "romantic restaurant with outdoor seating" — from exactly these structured details. An 80%-complete profile loses the qualified queries, which are the highest-intent ones.

### 4. Reviews: the text now matters as much as the stars

AI assistants summarize review *language* into your description. If your reviews say "fast, honest, fixed it same day," some version of that sentence becomes how ChatGPT introduces you. Practical implications: keep fresh reviews flowing (recency is weighted), respond to reviews (a sign of an active, real business — and your responses are also text the AI reads), and when you ask happy customers for reviews, ask them to mention *what* they loved — specific review text is raw material for your future AI description. Never fake or incentivize dishonestly; synthesized summaries amplify patterns, including suspicious ones.

### 5. Presence in the local lists AI already cites

Run the customer's question yourself on Perplexity and ChatGPT ("best [category] in [city]") and note the cited sources — typically local "best of" articles, city guides, and directory pages. Those specific pages are the electorate that keeps electing your competitors. Getting added to them — via outreach, genuinely earning inclusion, or claiming the directory profiles — moves your appearance rate faster than almost anything else, because you're joining sources already inside every answer.

### 6. A website AI can read: menus, services, and prices as text, not pictures

The local-business classic: the menu as a photo, the services list as a JPEG, the price sheet as a scanned PDF. Humans squint and cope; AI crawlers get nothing. Put your offerings in real HTML text, add LocalBusiness schema (and Menu/Service markup where it applies), and make sure the site works without JavaScript. When an assistant answers "does [your business] do [specific thing]?", the answer comes from text it could read — or it guesses, or names someone else.

## The 10-minute local self-check

Ask ChatGPT, Perplexity, and Gemini — a few times each, in fresh chats — the questions your customers ask: "best [category] in [city]," "[category] near [landmark] open Sunday," "is [your business] good?" Tally: are you mentioned, how are you described, and which sources get cited? The full method with prompts and scoring is in our companion guide: [How to ask ChatGPT about your own business](https://see-geo.com/blog/check-what-chatgpt-says-about-your-business). For local businesses the result is usually immediately diagnostic — you'll either see your review language reflected back at you (good sign) or discover the AI is working from a picture of your business that's years out of date.

## What this means for your weekly routine

Nothing exotic — a reweighting of work you already know: keep the profile current the way you'd keep the storefront clean, treat every review and response as future AI copy, check your listings for drift quarterly, and run the self-check with the seasons. The owners winning this channel early aren't doing secret AI tricks; they're doing ordinary local marketing with unusual consistency, aimed at machines as well as people.

And if you'd rather have the checking done for you: [SeeGeo's free audit](https://see-geo.com/#audit) covers the site-side essentials (crawler access, plain-language identity, schema, readable content) in about 30 seconds, and our tracking plans run the local prompt panel — "best [your category] in [your city]" and its variants, across the major assistants, month after month — so you see your share of recommendations as a trend instead of a quarterly surprise.

## Frequently asked questions

**Does ChatGPT use Google Business Profile data?**
Not through a direct feed — but AI assistants retrieve from the local-data ecosystem your profile anchors: directory listings, maps-derived data, review platforms, and pages that syndicate your business information. A complete, consistent profile improves the entire pool of sources AI draws from, which is why it remains the single highest-leverage local asset.

**How do I get my business recommended by ChatGPT for "near me" searches?**
Six levers: identical name/address/category everywhere; a homepage that plainly states what and where you are; a complete Google Business Profile treated as content; fresh, specific reviews (the text becomes your AI description); presence in the local "best of" pages AI already cites; and a website whose offerings are readable text with LocalBusiness schema, not images.

**Do online reviews affect AI recommendations?**
Strongly — and the text matters as much as the star rating. Assistants synthesize review language into how they describe and justify recommending you. Fresh, detailed reviews mentioning specific services give the AI accurate raw material; a thin or stale review footprint leaves it guessing or defaulting to competitors with richer signals.

**My business shows up on Google Maps but not in ChatGPT. Why?**
Maps ranking and AI recommendations draw on overlapping but different sources. Common causes: your website doesn't clearly state what you are (so the AI can't confidently classify you), you're absent from the "best of" articles and directories AI answers cite, or your site's content isn't readable without JavaScript. Run the [self-check](https://see-geo.com/blog/check-what-chatgpt-says-about-your-business) to see which pattern fits, then work that specific lever.

**Is AI local search big enough to matter for a small business yet?**
Yes — and it crossed that line recently. [BrightLocal's 2026 survey](https://www.brightlocal.com/research/lcrs-ai-trust/) found 45% of consumers now use AI for local business recommendations, up from 6% a year earlier — the fastest channel shift in local discovery since smartphones. The volume varies by category, but the trajectory is one-directional, and the work that wins it also improves your Google presence — it pays twice.

---

# How AI-visible is your website platform by default? What the docs, the research, and the live files actually say
Source: https://see-geo.com/blog/website-platform-ai-visibility-defaults · Updated 2026-08-18

> No major website platform blocks AI crawlers by default — and the 'block AI' toggles don't do what owners think. Docs + live-file checks, cited.

**Quick answer: none of the six major small-business website platforms — Shopify, Wix, Squarespace, WordPress (.com and self-hosted), Webflow, GoDaddy Website Builder — blocks a single AI crawler by default. We verified that two ways: by reading every platform's own documentation, and by checking the robots.txt of 45 live sites across four of the platforms (zero blocked GPTBot). The differences that actually matter sit elsewhere: two platforms now generate an llms.txt for you without asking (Shopify and Wix), the "block AI" toggles that Squarespace and WordPress.com offer quietly exempt the crawlers that power AI *search* (so flipping them doesn't remove you from ChatGPT or Perplexity answers), Webflow won't let you edit robots.txt at all without a paid plan, and GoDaddy documents no AI controls whatsoever.** This is a research synthesis with light verification, not a lab test — every claim below is cited to a primary source with its retrieval date, and where the evidence is practitioner-grade rather than official, we say so. As far as we can tell, nobody has published a controlled head-to-head platform comparison; we name that gap honestly at the end.

**Jump to your platform:** [Shopify](#shopify) · [WordPress](#wordpress-com-and-self-hosted) · [Wix](#wix) · [Squarespace](#squarespace) · [Webflow](#webflow) · [GoDaddy](#godaddy-website-builder)

## First, the finding that reframes the question

The scary story — "your website platform is silently blocking AI" — is not what the evidence shows. Large-scale measurement finds AI-crawler blocking concentrated among big publishers: a UC San Diego study presented at the Internet Measurement Conference 2025 found 12–14% of the web's top 5,000 sites fully disallow at least one AI crawler ([Liu et al., IMC 2025](https://arxiv.org/abs/2411.15091)), and the blocking wave concentrates in publishing — the Reuters Institute counted 48% of top news sites blocking OpenAI's crawlers by the end of 2023 ([Fletcher, 2024](https://reutersinstitute.politics.ox.ac.uk/how-many-news-websites-block-ai-crawlers)). Small-business sites behave differently: [our own 107-site audit](https://see-geo.com/blog/how-many-websites-block-ai-crawlers) found only 2.8% blocking any major AI crawler.

The platform layer is consistent with that. All six platforms ship permissive robots.txt files, per their own documentation — and our spot-check of 45 live platform sites found **zero** blocking GPTBot.

The documented technical risk for AI visibility is different: **rendering**. Vercel and MERJ measured how AI crawlers actually behave and found that none of the major ones execute JavaScript — OpenAI's, Anthropic's, Meta's, and Perplexity's crawlers all fetch JavaScript files but never run them; Google's Gemini (riding Googlebot's infrastructure) is the main exception ([Vercel, Dec 2024](https://vercel.com/blog/the-rise-of-the-ai-crawler)). A site whose content only exists after JavaScript runs is invisible to most AI systems.

Here's the twist for platform users: **that risk mostly spares you.** Wix states outright that its infrastructure is server-side rendered ([Wix SEO features](https://www.wix.com/seo/features)); Webflow publishes compiled static HTML (practitioner-documented — Webflow's own docs don't state the rendering model); Squarespace serves core content server-side (practitioner-observed, same caveat); and the GoDaddy sites we fetched carry their full text in the raw HTML. The JavaScript-invisibility problem concentrates in custom-built sites, not platform sites. If you're on a major platform, your AI-visibility problems are almost certainly [content and clarity problems](https://see-geo.com/blog/how-many-websites-block-ai-crawlers), not plumbing.

## The comparison at a glance

| Platform | AI bots blocked by default | "Block AI" setting | llms.txt | robots.txt editable |
|---|---|---|---|---|
| Shopify | None | No toggle — theme template edit only | **Auto-generated** (since May 2026) | Via `robots.txt.liquid` (unsupported customization) |
| WordPress.com | None | Opt-out toggle, off by default — **doesn't block OAI-SearchBot** | No | No (platform-managed) |
| WordPress self-hosted | None | Nothing in core; SEO plugins offer manual tools | Via plugins (AIOSEO: on by default; Yoast/Rank Math: opt-in) | Yes |
| Wix | None | No toggle — manual robots.txt edit | **Auto-generated** (premium + custom domain) | Yes |
| Squarespace | None | Toggle, off by default — **doesn't block PerplexityBot, OAI-SearchBot, or ChatGPT-User** | Manual, disabled by default | No (platform-managed) |
| Webflow | None | No toggle — and robots.txt editing **requires a paid plan** | Manual upload, custom domains only | Paid plans only |
| GoDaddy Website Builder | None (observed) | None documented | None documented | Not documented |

*Sources for every cell are linked in the platform sections below. Retrieved 2026-08-18; platforms change defaults silently — see the update policy at the end.*

## The toggle finding: training opt-outs dressed as visibility opt-outs

Two platforms offer a "block AI" checkbox. Both lists have the same hole, and it's the most consequential detail in this report.

**Squarespace's** "Block known artificial intelligence crawlers" setting (off by default) writes robots.txt rules for 26 bots — GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and more. Not on the list: **PerplexityBot, OAI-SearchBot, and ChatGPT-User** ([Squarespace help, retrieved 2026-08-18](https://support.squarespace.com/hc/en-us/articles/360022347072-Request-that-AI-models-exclude-your-site)).

**WordPress.com's** "Prevent third-party sharing" setting (also off by default) appends blocks for thirteen bots including GPTBot, ClaudeBot, and PerplexityBot — but not **OAI-SearchBot** ([WordPress.com support](https://wordpress.com/support/privacy-settings/make-your-website-public/), confirmed against a live opted-out site's robots.txt).

Why that matters: OpenAI and Perplexity split their crawlers by purpose — GPTBot gathers training data, while OAI-SearchBot and ChatGPT-User power live search and citations. (Most platform guides miss this distinction entirely.) So an owner who flips either toggle has opted out of model *training* — but their site still appears in ChatGPT search results and Perplexity answers. If you flipped the toggle to stay visible while limiting training use, that's arguably exactly what you wanted — but nothing in either interface tells you that's the deal you're getting.

## llms.txt became a platform decision while nobody was looking

In [our 107-site study](https://see-geo.com/blog/how-many-websites-block-ai-crawlers), 25.2% of small-business sites had an llms.txt file — and almost none of the owners put it there. The platform docs now confirm the mechanism:

- **Shopify** serves `/llms.txt` (mirroring a primary `/agents.md`) on every store, customizable via theme templates ([Shopify changelog, May 2026](https://shopify.dev/changelog/customize-llmstxt-llms-fulltxt-and-agentsmd)).
- **Wix** "automatically generates and maintains" llms.txt for premium sites with custom domains, with an opt-out toggle ([Wix help](https://support.wix.com/en/article/understanding-your-sites-llmstxt-file)). Our spot-check found 6 of 7 live Wix sites serving a ~3KB llms.txt.
- **All in One SEO**, one of WordPress's biggest SEO plugins, generates llms.txt **by default** once installed ([AIOSEO docs](https://aioseo.com/docs/how-to-create-an-llms-txt-using-all-in-one-seo/)); Yoast and Rank Math offer it opt-in.
- **Squarespace** ([manual, disabled by default](https://support.squarespace.com/hc/en-us/articles/47434125611277-Create-an-llms-txt-file)) and **Webflow** ([manual upload](https://help.webflow.com/hc/en-us/articles/43240104183315-Upload-an-llms-txt-file-to-your-site), which candidly calls the format "experimental") leave it to the owner. GoDaddy doesn't mention it.

Whether llms.txt matters is a separate question — [no major AI company has committed to reading it](https://see-geo.com/blog/ai-visibility-glossary) — but adoption statistics now measure platform product decisions, not site-owner behavior. Any study (including ours) that counts llms.txt files is partly counting Shopify and Wix market share.

## Shopify

Everything is allowed by default and bot management happens at Shopify's network layer — their docs say stores are readable by "search engines and large language models" with no action needed ([crawling docs](https://help.shopify.com/en/manual/promoting-marketing/seo/crawling-your-store)). There is **no admin toggle for AI crawlers** — posts describing one are talking about merchants' own Cloudflare accounts. Control means editing a theme template ([robots.txt.liquid](https://shopify.dev/docs/storefronts/themes/architecture/templates/robots-txt-liquid)), which Shopify calls an unsupported customization. The subtle one: **robots.txt only guards one of two doors.** Shopify Catalog syndicates eligible product data to AI shopping surfaces (ChatGPT, Microsoft Copilot) by default, and blocking crawlers doesn't stop it — that has [its own per-platform opt-outs](https://help.shopify.com/en/manual/online-sales-channels/agentic-storefronts/products). A practitioner failure mode worth checking on older stores: a 2023-era "block all AI bots" snippet copy-pasted into robots.txt.liquid and forgotten.

## WordPress (.com and self-hosted)

Core WordPress ships a two-line robots.txt that blocks nothing but the admin area ([developer reference](https://developer.wordpress.org/reference/functions/do_robots/)), and none of the three big SEO plugins (Yoast, AIOSEO, Rank Math) blocks any AI bot by default. WordPress.com adds the opt-out toggle described above — off by default, and porous to OpenAI's search crawler when on. Self-hosted owners hold full control and full responsibility: the same 2023 blocklist snippet circulates widely in WordPress tutorials ([the canonical version](https://neil-clarke.com/block-the-bots-that-feed-ai-models/) dates to August 2023), and a copied-then-forgotten robots.txt is the likeliest way a small business is *accidentally* invisible today.

## Wix

Permissive by default with a real robots.txt editor, no AI toggle ([blocking is a documented manual edit](https://support.wix.com/en/article/blocking-ai-crawlers-from-your-site)), first-party SSR — and the auto-generated llms.txt above. Of the six, Wix currently ships the most AI-visibility infrastructure without being asked.

## Squarespace

Platform-managed robots.txt you can't edit directly, with the off-by-default toggle — and one detail our live-file check surfaced: Squarespace's generated robots.txt **names the AI-bot roster explicitly on every site we checked** (15/15), with path-level rules that don't block content. The scaffolding is pre-built; the toggle just flips it. Remember what the toggle doesn't cover (Perplexity, OpenAI search) before treating it as an off switch.

## Webflow

The default robots.txt contains only a sitemap line — nothing blocked. The catch is access: adding robots.txt rules [requires a paid Site plan](https://help.webflow.com/hc/en-us/articles/41954183592851-Request-to-block-specific-bots-from-visiting-your-site), making Webflow the only platform here where *blocking* is a premium feature. Webflow also documents a Content-Signal HTTP header for granular AI preferences (train/search/input) — new, and its real-world effect depends entirely on crawler cooperation.

## GoDaddy Website Builder

The least documented surface of the six: no robots.txt documentation, no AI setting, no llms.txt for Websites + Marketing sites (their [SEO help article](https://www.godaddy.com/help/improve-my-websites-seo-20110) mentions none of it). Live sites we fetched serve a near-empty robots.txt and server-rendered HTML — fine defaults, zero control. Watch item: GoDaddy announced a Cloudflare partnership (April 2026) to integrate AI Crawl Control into its hosting platform; whether it reaches Website Builder, and with what defaults, is unannounced.

## What we checked ourselves, and the gap we're naming

The documentation claims above were spot-verified against reality: we fetched robots.txt and llms.txt (public files only, honest user-agent, rate-limited, aggregate reporting) from 45 live sites fingerprinted as Shopify (8), WordPress (15), Wix (7), and Squarespace (15). Results: zero sites blocked GPTBot; every Squarespace robots.txt named the AI-bot roster; 8/8 Shopify and 6/7 Wix sites served llms.txt versus 1/15 WordPress and 0/15 Squarespace — the platform-decision finding, visible in the wild. (Webflow and GoDaddy verification is pending a large enough site sample; their sections above rest on documentation and small-n observation.)

What nobody has published — us included — is a **controlled head-to-head test**: the same business, same content, built on each platform, audited identically. Everything above is synthesis of documentation plus observation of sites that differ in a thousand ways besides their platform. We checked for such a study and couldn't find one; if it exists, we'll link it here. If it continues not to exist, running it is on our roadmap.

One more honesty note, borrowed from a practitioner who re-tested his own AI-crawler findings six months later and found [11 of 17 had changed](https://www.wislr.com/articles/ai-bot-behavior-log-analysis): this field rots fast. Cloudflare alone changed AI-crawler defaults for new domains in July 2025 and [again for September 2026](https://blog.cloudflare.com/content-independence-day-ai-options/). Every claim in this report carries its retrieval date (2026-08-18), and we re-verify quarterly — the changes become their own report.

## What to actually do, whatever your platform

Your platform's defaults are almost certainly not your problem — which means the fixes are yours to make, not theirs. Check the three things that [actually separated visible from invisible sites in our audit data](https://see-geo.com/blog/how-many-websites-block-ai-crawlers): whether your homepage plainly says what you do and where, whether you have any structured data, and whether a stale blocklist snippet is sitting in your robots.txt. Or let [the free SeeGeo audit](https://see-geo.com/#audit) check all of it — including your actual, live robots.txt as the crawlers see it — in about 30 seconds.

## Frequently asked questions

**Does Shopify block ChatGPT or other AI crawlers?**
No. Shopify's default robots.txt blocks no AI crawlers, and Shopify states stores are readable by LLMs with no merchant action. There's no admin toggle either way — blocking requires editing a theme template. Separately, Shopify Catalog shares eligible product data with AI shopping surfaces by default, governed by its own opt-outs, not robots.txt.

**If I turn on Squarespace's or WordPress.com's "block AI" setting, will I disappear from ChatGPT?**
Mostly no. Both toggles block training crawlers (GPTBot, ClaudeBot, CCBot, and others) but leave out the crawlers that power live AI search — OAI-SearchBot and ChatGPT-User on both platforms, plus PerplexityBot on Squarespace. Your content can still be retrieved and cited in AI answers with the toggle on.

**Which website builder is best for AI visibility?**
By defaults alone, none of them blocks AI, so "best" comes down to what's built for you versus locked from you: Wix and Shopify generate llms.txt automatically and serve crawler-readable HTML; Webflow and Squarespace are solid but leave AI files manual (and Webflow paywalls robots.txt control); GoDaddy offers no documented AI controls. No platform choice substitutes for clear content — that's where small-business sites actually fail.

**Do I need an llms.txt file?**
It's cheap insurance, not a ranking lever — no major AI company has committed to reading it. If your platform generates one (Shopify, Wix, or WordPress with AIOSEO), check what's in it rather than adding one. The interesting fact is that adoption statistics now mostly measure platform defaults, not owner intent.

**How current is this report?**
All sources and live-file checks are dated 2026-08-18, and each claim links its source. Platforms change defaults silently — we re-verify quarterly and publish what changed.

---

# The AI visibility glossary: 26 terms business owners keep seeing, explained in plain English
Source: https://see-geo.com/blog/ai-visibility-glossary · Updated 2026-08-17

> GEO, AEO, llms.txt, query fan-out, GPTBot — AI search is drowning in jargon. 26 terms in plain English, each tagged established, emerging, or mostly hype.

**Quick answer: this glossary defines the 26 terms you'll actually encounter when learning how AI systems like ChatGPT, Perplexity, and Google's AI find and recommend businesses — in plain English, with each term tagged honestly: Established (proven, act on it), Emerging (real but young, watch it), or Mostly Hype (skip it, and we say why).** The tags matter because this field is two years old and full of confident-sounding jargon selling uncertain things. Bookmark this page; we update it quarterly.

---

## AEO (Answer Engine Optimization)

**Established (as a term).** The practice of optimizing content to appear in AI-generated answers. Functionally the same thing as GEO — different communities coined different acronyms for the same work. If you see AEO, GEO, or "AI visibility," assume they mean the same discipline.

## Agentic commerce

**Emerging.** The next phase after AI recommendations: AI agents that don't just suggest businesses but transact with them — browsing, booking, and buying on a user's behalf. Payment and checkout protocols from major platforms emerged through 2025–26, but consumer adoption is early. Worth watching, not yet worth re-architecting your site for.

## AI crawler

**Established.** A bot an AI company uses to read websites, the way Googlebot reads them for search. Each company runs its own: GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), Google-Extended (Google's AI training). Your site can allow or block each one individually — and many sites block them without their owners knowing. Full article: [Is your website visible to ChatGPT?](https://see-geo.com/blog/is-your-website-visible-to-chatgpt)

## AI Overviews

**Established.** The AI-generated summary Google shows above traditional search results for many queries. Sources cited in AI Overviews draw heavily from pages that already rank well in normal Google search — meaning your traditional SEO substantially feeds this surface.

## AI Mode

**Established.** Google's fuller conversational search experience (a step beyond AI Overviews), which answers questions through dialogue and runs aggressive query fan-out behind the scenes. See *query fan-out*.

## Citation

**Established.** When an AI answer links to or names a specific source for a claim — the AI-era equivalent of ranking. Citations are the measurable win in AI visibility: they're how a reader discovers your business from inside an answer. Compare *mention*.

## ClaudeBot

**Established.** Anthropic's web crawler, which gathers content that informs Claude. One of the user-agents to check in your robots.txt — our [crawler library](https://see-geo.com/bots) documents each one. Full article: [Is your website visible to ChatGPT?](https://see-geo.com/blog/is-your-website-visible-to-chatgpt)

## Crawl-to-referral ratio

**Emerging.** How many times AI crawlers read your site versus how many human visitors AI systems send back. Published CDN data shows crawls vastly outnumber referred clicks — useful context for deciding how much AI traffic to expect, and a reason the "visibility" benefit of AI is partly brand exposure rather than clicks.

## Entity

**Established.** What AI and search systems understand your business to *be* — its name, category, location, and how confidently the system can classify it. Inconsistent descriptions across the web create a blurry entity, and blurry entities don't get recommended. The fix is unglamorous consistency: same name, same category description, everywhere.

## Extractability

**Established.** How easily an AI can lift a usable, self-contained answer from your page. Answer-first openings, question-shaped headings, short standalone passages, and tables all raise it. The single most within-your-control factor in AI visibility. Full article: [The GEO writing checklist](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt)

## GEO (Generative Engine Optimization)

**Established (as a term).** The [umbrella practice](https://see-geo.com/blog/what-is-geo) of improving how often and how favorably AI systems mention, cite, or recommend your business. Coined in a 2023 Princeton research paper. Not to be confused with "geo" as in geographic targeting — an entirely different, older marketing term. Full article: [SEO vs GEO](https://see-geo.com/blog/seo-vs-geo-difference)

## GPTBot

**Established.** OpenAI's main web crawler. Blocking it in robots.txt (or via CDN settings) cuts your site off from the systems behind ChatGPT — one of the most common critical findings in our audits, and often switched on by a CDN default the owner never chose. Full article: [Should you block AI crawlers?](https://see-geo.com/blog/should-you-block-ai-crawlers)

## Grounding

**Established.** When an AI system bases its answer on retrieved, current sources (with citations) rather than only on what it memorized in training. Grounded answers are why fresh, extractable content can win citations quickly — the AI is actively looking for sources at answer time.

## Hallucination

**Established.** When an AI states something false with confidence. Relevant to businesses because AI systems can describe you inaccurately; the practical defense is consistent, clear information about your business everywhere the AI might learn from. Checking what AI says about you should be routine, like Googling yourself.

## llms.txt

**Emerging, contested.** A proposed standard: a markdown file at your site root giving AI systems a curated index of your key content. Cheap to add, no downside — but Google has said on the record it ignores the file, and no major engine has committed to honoring it. A 20-minute lottery ticket, not a strategy. Full article: [llms.txt — an honest verdict](https://see-geo.com/blog/llms-txt-honest-verdict)

## Mention

**Established.** When your brand is named anywhere — in an AI answer or on any webpage — with or without a link. AI systems learn about brands from mentions across the whole web, which is why a Reddit thread or a "best of" listicle that names you (no link needed) still builds your AI visibility. The AI-era loosening of the old backlink obsession.

## MCP (Model Context Protocol)

**Emerging.** A technical standard letting AI assistants connect directly to tools and data sources. For software companies, "being a tool the AI can call" is a real distribution strategy; for typical small businesses, it's not yet relevant. Watch, don't build.

## Non-determinism

**Established.** The same question asked twice produces different AI answers. Run any category prompt through the same engine a few times and compare the brand lists — they rarely match exactly, which you can verify yourself in minutes. The single most important measurement fact in this field: it's why any tool selling you a fixed "AI rank" is selling snake oil, and why honest measurement means percentages across many samples. See *share of voice*.

## PerplexityBot

**Established.** Perplexity's crawler, feeding its own search index. Perplexity is the most retrieval-driven major engine — it can cite a page published days ago — making it the fastest feedback loop for new content.

## Prompt panel

**Emerging (as standard practice).** A fixed set of realistic customer questions ("best [category] in [city]," "alternatives to [competitor]") run repeatedly across AI engines to measure whether and how your business appears. The AI-era equivalent of rank tracking — with repeated sampling required because of *non-determinism*.

## Query fan-out

**Established (behavior); Emerging (optimization practice).** When an AI system silently breaks one question into many sub-queries, searches them all, and assembles the answer from whichever sources best answer each piece. The practical upshot: cover the cluster of sub-questions around your topic, not just the head keyword. Full article: [Query fan-out, explained simply](https://see-geo.com/blog/query-fan-out-explained)

## Retrieval

**Established.** The step where an AI system fetches relevant content from an index before writing its answer. Different engines retrieve from different indexes (ChatGPT leans on Bing's, Gemini on Google's, Perplexity on its own) — which is why fixing your Bing presence can fix your ChatGPT visibility, a connection most tools miss.

## Schema (structured data)

**Established.** Machine-readable labels (usually JSON-LD) describing what a page contains — a product, a FAQ, a local business. The closest thing to a shared language between your site and both search and AI systems. FAQ and HowTo schema matter most for AI visibility because they package content in the question-answer shape AI extracts from.

## Share of voice

**Established (metric); measurement quality varies wildly.** The percentage of AI answers in your category that mention you versus competitors, measured across a prompt panel over time. The honest headline metric of AI visibility — as a sampled percentage with a margin of error, never a precise rank. Tools that report it from single runs are reporting noise.

## Vector/embedding optimization

**Mostly Hype.** Paid services claiming to tune your content for AI systems' internal semantic representations. No engine exposes its embedding space, the claims are unfalsifiable, and the legitimate part (write clearly about one topic per passage) is just good extractability wearing a lab coat. Save your money.

## Visibility score

**Emerging; read the methodology.** A single number summarizing how findable your business is to AI systems. Useful as a tracked trend, misleading as an absolute — the number is only as honest as the methodology behind it (sampling, engines covered, what's weighted). Any score without a published methodology is marketing.

---

## How to use this glossary

If you're new: read *AI crawler*, *extractability*, *mention*, and *non-determinism* first — those four concepts are 80% of the practical picture. If you're evaluating tools or agencies: the tags are your filter — anyone selling certainty on an *Emerging* term, or anything at all on a *Mostly Hype* term, has told you what you need to know. For deeper definitions of the search-engine fundamentals behind these, our [site glossary](https://see-geo.com/glossary) goes term by term.

And if you want the applied version: [SeeGeo's free audit](https://see-geo.com/#audit) checks your site against the *Established* items on this list — crawler access, extractability, schema, entity clarity — in about 30 seconds, free, no signup for your score.

*Last updated: August 2026. This field moves fast; we revise tags and add terms quarterly. Spotted a term we should add? Tell us at info@see-geo.com.*

---

# We audited 107 small-business websites for AI visibility. The problem isn't what everyone says it is
Source: https://see-geo.com/blog/how-many-websites-block-ai-crawlers · Updated 2026-08-17

> We ran our audit engine over 107 real small-business sites in 7 countries. Only 2.8% block AI crawlers — the real gaps are structured data and clarity.

**Quick answer: we ran SeeGeo's audit engine over 107 real small-business websites — restaurants, trades, clinics, shops, professional services across the US, UK, Canada, Australia, Ireland, France, and Germany. Only 3 of 107 (2.8%) explicitly block any major AI crawler in robots.txt, so the popular "your site is blocking AI without you knowing" story is rare in this sample. The real gaps are quieter: 29 of 107 (27.1%) have no structured data anywhere, 69 of 107 (64.5%) never plainly say what the business does at the top of the homepage, and 101 of 107 graded C or worse on overall AI visibility. One genuine surprise: 27 of 107 (25.2%) already have an llms.txt file — and almost none of them put it there.** Every number below comes from the actual runs, with the full methodology published at the end of this post.

## How many small-business sites block AI crawlers?

3 of 107 (2.8%). Those three explicitly name AI crawlers in hand-written robots.txt groups and disallow them from the whole site. Counting wildcard rules that happen to catch AI bots too, the number stays 3 of 107 — no site in this sample is root-blocked for an AI crawler by accident of a `User-agent: *` rule.

That number is worth sitting with, because the industry narrative — ours included, at times — leans on the idea that small businesses are accidentally invisible to AI because something is blocking the crawlers. In this sample, that's the exception. Nearly everyone lets AI in. What happens after the crawler gets in is where the visibility is lost.

## Which AI crawler is blocked most?

| Crawler | Explicit robots.txt block | Any root block |
|---|---|---|
| GPTBot (OpenAI, training) | 3 of 107 (2.8%) | 3 of 107 (2.8%) |
| ClaudeBot (Anthropic) | 3 of 107 (2.8%) | 3 of 107 (2.8%) |
| Google-Extended (Gemini training) | 2 of 107 (1.9%) | 2 of 107 (1.9%) |
| CCBot (Common Crawl) | 2 of 107 (1.9%) | 2 of 107 (1.9%) |
| PerplexityBot | 1 of 107 (0.9%) | 1 of 107 (0.9%) |
| OAI-SearchBot (ChatGPT citations) | 0 of 107 (0.0%) | 0 of 107 (0.0%) |
| Googlebot | 0 of 107 (0.0%) | 0 of 107 (0.0%) |
| Bingbot | 0 of 107 (0.0%) | 0 of 107 (0.0%) |

GPTBot and ClaudeBot tie for most-blocked at 3 sites each. Notably, OAI-SearchBot — the crawler that decides whether ChatGPT can *cite* a site — is blocked by nobody in this sample.

## Who blocked them, and was it deliberate?

Among the 3 blocking sites: 2 show hand-written robots.txt rules with no CDN or firewall signals, and 1 pairs its robots.txt rules with a Cloudflare front. All three name the AI bots explicitly in their own robots.txt groups — these look like decisions, not defaults. 0 of 107 (0.0%) sites carry the signature of Cloudflare's managed robots.txt in this sample.

One honest caveat that cuts the other way: 31 of 107 (29.0%) of the audited sites sit behind Cloudflare, and CDN-level bot blocking that leaves no external signature is not measurable from outside. Our robots.txt numbers are exact; our CDN numbers are a floor. (A related receipt from our own run: one site served HTTP 403 to our honestly-identified audit crawler while serving a normal page to a browser — network-level bot filtering is real, which is why sites that blocked *our* checker were excluded from every statistic rather than counted as "blocks AI.")

## The JavaScript problem is smaller than advertised

3 of 107 (2.8%) sites serve a homepage whose primary content doesn't exist unless JavaScript runs — which is how most AI crawlers read the web, so those three are effectively blank to AI. We expected more. Modern site builders appear to server-render well enough that the "your beautiful JS site is invisible" warning, while real, is also the exception in this sample.

## Structured data is where the sample actually fails

- No JSON-LD structured data anywhere on the crawled pages: **29 of 107 (27.1%)**
- JSON-LD on the homepage: 75 of 107 (70.1%)
- JSON-LD beyond the homepage: 69 of 107 (64.5%)
- A plain-language "what we do" statement near the top of the homepage: only **38 of 107 (35.5%)** — meaning 69 of 107 (64.5%) make both humans and machines infer what the business is
- Overall grades from the audit engine: B 6 · C 38 · D 38 · F 25 — 101 of 107 (94.4%) graded C or worse

This is the actual shape of the small-business AI-visibility problem in our data: the door is open, but the room is unlabeled.

## The llms.txt surprise: adoption by default

27 of 107 (25.2%) sites have an llms.txt file — a number that would suggest remarkable adoption of a contested young standard. It doesn't. We re-fetched and classified every one:

- Every Shopify store in the sample (8 of 8) carries Shopify's auto-generated "Agent Instructions" llms.txt.
- Every Wix hit carries Wix's auto-generated markdown file.
- Several WordPress hits are generated by the All in One SEO plugin, which stamps its own signature into the file.

In other words: **platforms are adopting llms.txt on their customers' behalf.** Most of these owners almost certainly don't know the file exists. Whether the standard matters is a separate question — [our honest verdict on llms.txt](https://see-geo.com/blog/llms-txt-honest-verdict) is unchanged — but "who is adopting llms.txt" now has a data-backed answer: software vendors, not site owners.

## What should a business owner do?

The study's practical order of operations, given what actually failed:

1. **Say what you do, in plain words, at the top of your homepage.** The single most-failed check (64.5%) and the cheapest fix on this list.
2. **Add structured data** — at minimum an Organization or LocalBusiness block, plus FAQ markup where you genuinely answer questions. A quarter of the sample has none at all.
3. **Only then worry about crawler access.** Check it — it takes seconds and the rare failure is severe — but in 97% of cases in this sample, the door was already open.

[SeeGeo's free audit](https://see-geo.com/#audit) runs all of these checks on your site in about 30 seconds, no signup. Deeper dives: [should you block AI crawlers?](https://see-geo.com/blog/should-you-block-ai-crawlers), [is your website visible to ChatGPT?](https://see-geo.com/blog/is-your-website-visible-to-chatgpt), and [the AI visibility glossary](https://see-geo.com/blog/ai-visibility-glossary).

## Frequently asked questions

**Does blocking GPTBot remove a site from ChatGPT?**
Blocking GPTBot stops OpenAI's model-training crawler. ChatGPT's search citations depend on OAI-SearchBot — blocked by 0 of 107 sites here. Blocking both makes citation practically impossible; blocking only GPTBot mostly affects what future models learn.

**Is robots.txt the only way sites block AI crawlers?**
No. CDNs and firewalls can block at the network level with no robots.txt trace. We report detected CDN signals separately (32 of 107 sites show a WAF/CDN signature), and we treat our CDN numbers as a floor because signature-free blocking is invisible from outside.

**Did the three blocking sites choose to block AI?**
It looks that way: all three name AI bots explicitly in hand-written robots.txt groups, and none carries the signature of a CDN-managed robots.txt. In this sample, blocking was rare but deliberate.

**Why is your llms.txt number so much higher than reported adoption?**
Because platforms ship it by default now. Shopify auto-generates an llms.txt for every store, Wix generates one, and popular WordPress SEO plugins emit one. Owner-initiated adoption in this sample is a small fraction of the 25.2%.

**Can I see the data?**
Yes — the aggregate data, findings, methodology, and the llms.txt classification are published on GitHub: [github.com/Tunisian-Aaron/ai-visibility-study-2026](https://github.com/Tunisian-Aaron/ai-visibility-study-2026) (CC BY 4.0). The domain list and per-site results are deliberately not published — this study is statistics, not a wall of shame.

## Full methodology

*Published verbatim; runs conducted 2026-08-17.*

**Sample.** 120 candidate domains were collected from public "best of" listicles, local directory pages, and trade-association member lists across 12 business categories (restaurants/cafés, trades, clinics, salons/fitness, independent retail, professional services, auto repair, artisan producers) and 7 countries (US, UK, Ireland, Canada, Australia, France, Germany). Selection rules: small businesses only; national chains, franchises, and enterprise-scale companies excluded during collection (43 exclusions logged with reasons); aggregator/directory domains excluded; the business's own site only; no site owned by SeeGeo or anyone we have a relationship with. This is a convenience sample — it is not random and not representative of the web. Findings describe this sample only, and we do not name the audited businesses.

**Runs.** Each site was crawled once on 2026-08-17 by the same engine that powers SeeGeo's public audit: homepage, up to ~8 site pages, robots.txt, sitemap, and an llms.txt probe. User-agent: `SeeGeoAudit/1.0 (+https://see-geo.com/bots)` — honestly identified, with an explanation page. Same-domain requests were spaced at least 1 second apart; one visit per site, with a single retry after network failure. 13 of 120 candidates were excluded at run time (8 unreachable, 3 domains moved, 2 blocked our checker); they appear in no denominator. A site that blocks or challenges our checker is recorded as "could not verify" — never counted as "blocks AI crawlers."

**Definitions.** *Blocks bot X (explicit)*: robots.txt contains a user-agent group explicitly matching X, and evaluating X against the site root yields disallowed; wildcard-only disallows are tallied separately and never merged into the explicit number. *CDN/WAF signals*: response-header signatures only (e.g. `server: cloudflare`); CDN blocking that leaves no signature is not externally measurable, so CDN numbers are a floor. *Primary content missing without JavaScript*: the engine's deterministic verdict that the crawled HTML — no JS execution, which is what most AI crawlers see — contains no usable primary content. *Structured data*: parseable JSON-LD; "beyond the homepage" means at least one non-homepage crawled page carries it. *llms.txt present*: HTTP 200 at /llms.txt with a non-HTML body; soft-404s serving the site shell count as absent, and one initial positive was corrected to absent on re-verification (logged). Every present file was re-fetched and classified by signature (Shopify auto-generated, Wix auto-generated, SEO-plugin output, other). *"What we do" statement*: the engine's shipped heuristic — English, French, and German patterns over the first 1,200 characters of homepage text plus the meta description.

**Limitations.** Convenience sample; results describe these 107 sites on this date, nothing more. CDN-level blocking is partially detectable at best. Bot rosters change; we evaluated the engine's roster as of the run date. Some "unreachable" exclusions may themselves be network-level blocks of our checker — indistinguishable from outside, which is one more reason excluded sites never enter a denominator. Subgroup statistics are reported only where n ≥ 20. The underlying aggregate data is published at [github.com/Tunisian-Aaron/ai-visibility-study-2026](https://github.com/Tunisian-Aaron/ai-visibility-study-2026).

---

# llms.txt: should your website have one? An honest verdict
Source: https://see-geo.com/blog/llms-txt-honest-verdict · Updated 2026-09-02

> Is llms.txt used by anyone, or a myth? Google says it does nothing; the industry recommends it anyway. What 152 real sites show, and a clear verdict.

**Quick answer: llms.txt is a proposed standard — a plain markdown file at your site's root (`yoursite.com/llms.txt`) that gives AI systems a curated, easy-to-read index of your most important content. Whether you need one is 2026's most genuinely contested SEO question: Google's Search team has said on the record that it does not use llms.txt and has no plans to, comparing it to the long-dead keywords meta tag — while parts of the industry recommend it as an emerging best practice. Our honest verdict: it's a 20-minute, zero-risk, low-probability bet — worth adding if you treat it as a lottery ticket, a waste of energy if you treat it as a strategy.**

Most coverage of llms.txt picks a side and oversells it. This article lays out what the file actually is, what each side's evidence says, and gives you a decision you can defend.

## What is llms.txt, exactly?

llms.txt is a proposal (originated by Jeremy Howard of Answer.AI in late 2024, documented at llmstxt.org) for a standardized file that solves a real problem: AI systems have small context windows and websites are bloated with navigation, scripts, and boilerplate that make extracting the substance hard. The file is written in markdown — human-readable, LLM-friendly — and contains a short description of the site plus a curated list of links to your most important pages, optionally annotated.

Think of it as the difference between handing someone your whole filing cabinet versus a one-page index of what matters. A companion variant, **llms-full.txt**, goes further and includes the full text of key content in one file.

**What llms.txt is *not*:** it's not robots.txt (which controls what crawlers *may* access — permission), and it's not a sitemap.xml (which lists *everything* for indexing — inventory). llms.txt is a *recommendation* — "here's what matters and where to find it." Three different files, three different jobs.

## What's the case against it?

The strongest evidence against comes from the biggest player. Google's Search Relations team has stated directly that Google does not support llms.txt — John Mueller publicly compared it to the keywords meta tag, the canonical example of a signal search engines learned to ignore — and Google's guidance states such files "will neither harm nor help" visibility because Google Search ignores them entirely. Since Google's AI surfaces (AI Overviews, AI Mode) draw on Google's index, llms.txt does nothing for the AI surface with the most users.

Beyond Google, the skeptic case has two more legs. First, **no major AI engine has publicly committed to honoring it** — OpenAI, Anthropic, and Perplexity have not announced that their answer pipelines use llms.txt. Second, practitioner server-log studies through 2025 found AI crawlers fetching llms.txt files inconsistently at best — being *fetched* occasionally is not the same as being *used* to shape answers. There is currently no rigorous public evidence that adding llms.txt increases citations or AI visibility.

## What's the case for it?

The advocate case is honest when it's framed as a bet, not a result:

**The problem it addresses is real.** AI systems genuinely do struggle with bloated HTML, and token efficiency genuinely matters to retrieval pipelines. A standard solving a real problem has a plausible adoption path — robots.txt itself started as an informal convention that everyone eventually honored.

**Adoption costs nothing and carries no penalty.** Even Google's dismissal concedes it "will neither harm nor help." The file takes minutes to create, doesn't interact with your rankings, and can't break anything.

**Some AI-adjacent tools and crawlers do read it.** Documentation platforms auto-generate it, some developer-focused AI tools consume it, and there are fetch records in server logs. Adoption is real among AI-forward software companies — it's engine *commitment* that's missing.

**If it wins, early adopters win retroactively.** Standards flip fast in this space. If any major engine announces support, sites with the file already in place benefit from day one — and everyone else spends that quarter catching up.

## What real websites actually do (our data, not opinion)

Two datasets we collected ourselves, both published with methodology:

- **107 small-business sites (August 2026 study):** 25.2% served an llms.txt file. Almost none of those were hand-written — the overwhelming majority were platform defaults, which is the tell: adoption is being driven by website builders shipping the file automatically, not by owners deciding it matters. [Full study →](https://see-geo.com/blog/how-many-websites-block-ai-crawlers)
- **45 live sites across six platforms (platform-defaults series):** Shopify serves an auto-generated llms.txt on every store we checked (8 of 8), Wix on most (6 of 7), and the rest don't. So whether your site "has" one is increasingly a fact about your platform, not a choice you made. [Platform defaults →](https://see-geo.com/blog/website-platform-ai-visibility-defaults)

The honest reading of both: a quarter of the web has the file, no engine has committed to reading it, and none of the sites that have it can show a citation they earned because of it. That's exactly what a low-probability bet looks like from the outside — neither a myth nor a strategy.

## So should you add one? Our verdict by situation

**If you have 20 spare minutes: yes, add it — with your expectations set correctly.** It's a costless hedge on a possible future, not a visibility tactic for the present. Rank it dead last on your GEO to-do list, after the things with actual evidence behind them (crawler access, answer-first content, statistics and citations, structured data, off-site mentions).

**If you're choosing where to spend limited effort: skip it without guilt.** Every hour spent hand-crafting an llms.txt is an hour not spent restructuring a page opening or getting into a source AI engines already cite — moves with measured effects (the KDD 2024 Princeton study's 30–40% visibility gains from statistics and citations). Google's own advice — put the effort into genuinely crawlable, well-structured content instead — is, on the evidence, correct about priorities even if you disagree about the file.

**If someone is charging you money to "implement llms.txt optimization": walk away.** A 20-minute markdown file sold as a paid deliverable, on a standard no engine honors, is the kind of certainty-theater this young industry has too much of.

## How to add llms.txt in 20 minutes (if you're taking the bet)

### Step 1: Create the file

**Create a plain text file** named `llms.txt`, in markdown, structured as: an H1 with your site/company name, a one-paragraph blockquote summary of what you do, then H2 sections (e.g., "Products," "Guides," "About") each containing a short list of links with one-line descriptions.

### Step 2: Curate ruthlessly

Pick 10–30 of your genuinely important pages, not your whole sitemap. The entire value proposition is curation — a dump defeats the purpose.

### Step 3: Upload it to your site root

It should resolve at `yoursite.com/llms.txt`, alongside robots.txt.

### Step 4: Keep it honest and current

Update it when your key pages change; a stale index is worse than none.

### Step 5: Optionally add llms-full.txt

If you have cornerstone content worth including in full, the companion file carries the full text. Most SMBs can skip this.

### Step 6: Re-check the landscape quarterly

The single fact that would change this article's verdict is a major engine announcing support. That's the trigger to move llms.txt from "lottery ticket" to "required".

## The bigger lesson: how to handle contested tactics in a young discipline

llms.txt is a case study in how to think about GEO advice generally. This field is roughly two years old; it has one landmark peer-reviewed study, a fast-moving practitioner consensus, and a strong commercial incentive to oversell every new tactic. The useful filter is a two-question test: *What's the cost if this does nothing?* and *What's the evidence it does something?* llms.txt scores near-zero cost, near-zero evidence — hence: fine as a hedge, foolish as a strategy. Plenty of louder tactics (paid "vector optimization," guaranteed "AI rankings") score high cost, zero evidence — those you run from.

That two-question honesty is also how we build [SeeGeo](https://see-geo.com): the audit flags as critical only what the evidence actually supports — blocked AI crawlers, content invisible without JavaScript, missing answer structure. llms.txt isn't on that list, and we won't sell you certainty about it in either direction. If a major engine ever commits to the standard, expect this article's verdict to change — in public, on this page.

---

## Frequently asked questions

**What is the difference between llms.txt and robots.txt?**
robots.txt controls permission — which crawlers may access which parts of your site. llms.txt is a recommendation — a curated, markdown-formatted index of your most important content for AI systems that choose to read it. robots.txt is universally honored; llms.txt is a proposal no major engine has committed to.

**Does Google use llms.txt?**
No. Google's Search Relations team has said on the record that Google does not support llms.txt and has no plans to, with John Mueller comparing it to the abandoned keywords meta tag. Google's guidance states the file neither helps nor harms, because Google ignores it.

**Does ChatGPT or Perplexity read llms.txt?**
Neither OpenAI nor Perplexity has publicly committed to using llms.txt in their answer pipelines. Server logs show AI crawlers occasionally fetching the file, but fetching is not the same as using it to shape answers. As of 2026 there is no rigorous public evidence llms.txt increases AI citations.

**Can llms.txt hurt my SEO?**
No. It's invisible to traditional search ranking systems and, per Google's own statement, neither helps nor harms. The only real cost is the time spent creating and maintaining it — and stale, inaccurate llms.txt files are the one way to make it counterproductive if an engine ever does read it.

**What should I do instead of (or before) llms.txt?**
The evidence-backed priority list: ensure AI crawlers aren't blocked (robots.txt and CDN settings), make key content readable without JavaScript, restructure important pages answer-first with statistics and cited sources, add FAQ/HowTo structured data, and build presence in the third-party sources AI engines already cite. Every one of those has stronger evidence than llms.txt; do them first.

**How do I create an llms.txt file?**
Write a markdown file: H1 with your company name, a blockquote paragraph summarizing what you do, then H2 sections listing 10–30 of your most important pages as links with one-line descriptions. Save it as llms.txt at your site root. Twenty minutes, curated rather than complete, updated when your key pages change.

**Does an llms.txt file do anything for AI visibility, or is that a myth?**
Today, measurably nothing you can point to: Google says it ignores the file, and neither OpenAI nor Perplexity has committed to using it. It isn't a myth that the file exists or that some crawlers fetch it — it's a myth that adding one has been shown to change AI answers. Treat it as a cheap bet on a standard that might be adopted later, not as a visibility tactic.

**Is llms.txt used by anyone?**
Some AI documentation tools and a few developer-facing crawlers read it, and server logs show occasional fetches from AI bots. What's missing is any major answer engine — Google, ChatGPT, Perplexity, Claude — confirming it shapes their answers. "Fetched sometimes" is the accurate status; "used to rank or cite" is not, as of 2026.

**Is llms.txt worth it? Is it necessary?**
Necessary: no — nothing you're doing for AI visibility depends on it. Worth it: only as a 20-minute, zero-risk lottery ticket after the evidence-backed work is done (crawler access, no-JavaScript readability, answer-first structure, structured data). If you're on Shopify or Wix you likely already have one and did nothing to get it.

---

# Query fan-out, explained simply: why AI answers 20 questions when you ask one
Source: https://see-geo.com/blog/query-fan-out-explained · Updated 2026-09-02

> Query fan-out (or fan out, fanout): how AI search splits one question into hidden sub-queries. What it is, a 10-minute no-tool analysis, how to optimize.

**Quick answer: query fan-out is how AI search systems — Google's AI Mode, ChatGPT, Perplexity — handle a question: instead of running one search for what the user typed, they silently break it into many smaller sub-queries ("fan it out"), search each one in parallel, and assemble the answer from whichever sources best answer each piece.** The practical consequence for your business: you're no longer competing to rank for one keyword — you're competing to be the best answer to a dozen hidden sub-questions you never see. Pages that cover those sub-questions are dramatically more likely to be cited: one 2025 correlation analysis found pages ranking for fan-out sub-queries were 161% more likely to be cited in AI Overviews than pages ranking only for the main query.

Here's how it works, how to see the hidden sub-queries yourself, and how to restructure content for it — in plain language.

## What actually happens when someone asks an AI a question?

Say a customer asks: *"What's the best running shoe for marathon training?"*

A traditional search engine would run that one query and rank pages against it. An AI search system instead **deconstructs** it into sub-queries — something like: "running shoes for long distances," "marathon shoe cushioning," "durable shoes for high mileage," "running shoes for narrow feet," "best running shoes 2026." It then **retrieves** results for each sub-query in parallel, **aggregates** the strongest sources across all of them, and **synthesizes** one answer — citing the handful of pages that best answered specific pieces.

The intensity varies by engine: analyses of engine behavior find Google's AI Mode runs the most aggressive fan-out (sometimes dozens of sub-queries), ChatGPT is moderate, and Perplexity stays comparatively focused. But all three work this way — which is why fan-out has been called the biggest shift in how search retrieves content since semantic search.

## Why should a business owner care?

Three reasons, each with a number attached:

**1. Sub-query coverage — not head-keyword ranking — predicts citations.** The ALM Corp analysis mentioned above (2025, Spearman correlation against Semrush data) found the 161% citation advantage for pages ranking on fan-out sub-queries — and, more striking, pages ranking *only* for sub-queries (not the main term at all) were still 49% more likely to earn citations than pages ranking only for the head term. Being the best answer to a fragment beats being a mediocre answer to the whole thing.

**2. You can rank #4 and still win the answer.** Roughly 52% of sources cited in Google AI Overviews rank somewhere in the top 10 — not necessarily #1 (AIOSEO, 2025). AI systems cite whoever answers the *sub-question* best, which means a smaller site with a sharper specific answer regularly gets cited over the #1 result. Fan-out is, quietly, the most small-business-friendly mechanic in modern search.

**3. Your rank tracker can't see any of this.** A traditional rank tracker measures your position for the head keyword. It has no idea whether you're being retrieved for the twelve invisible sub-queries — which is where the citation decision actually happens. If your high-ranking pages aren't showing up in AI answers, missing sub-query coverage is the most common explanation.

## How can I see the hidden sub-queries myself?

This is the part most guides skip, and it's genuinely useful — the fan-out isn't fully secret:

- **Google AI Mode** sometimes shows a "searches it ran" indicator on responses — read it; that's the literal fan-out list.
- **ChatGPT**: its search behavior can be observed in the browser's developer tools, and free browser extensions exist that extract the sub-queries from a ChatGPT response.
- **Gemini's Grounding API** exposes the queries it grounds against, for the technically inclined.
- **Low-tech proxies**: Google's "People Also Ask" boxes are typically a subset of fan-out sub-queries, and question-mapping tools (like AlsoAsked) chart the related-question tree around any topic.

Run your most important customer question through two of these and write down the sub-queries. That list is your content plan.

## How do I optimize for query fan-out? (Without falling in the trap)

The core move: **make each important page answer the cluster, not just the keyword.**

1. **Map the sub-query spectrum first.** For your money question, collect sub-queries from the methods above. Look especially for the two types analysts find most commonly missed: *comparative* sub-queries ("X vs Y," "alternatives to X") and *implicit* ones (the question behind the question — someone asking about marathon shoes implicitly asks about durability, price, and fit).

2. **Give each sub-query its own H2/H3 section, answered in the first one or two sentences.** Modular, passage-level structure matters because the AI selects and cites *sections*, not whole pages. (This is the same extractability principle as [answer-first writing](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt) — fan-out is *why* it works.)

3. **Front-load the page.** One analysis found 44.2% of all LLM citations come from the first 30% of a document. Your most valuable answers should not live in the basement of a 4,000-word page.

4. **Use tables and lists for the comparative sub-queries.** Structured formats get extracted for comparison-type sub-queries far more reliably than prose.

5. **One deep page beats ten thin ones.** Fan-out rewards comprehensive topical coverage on a single resilient page — the opposite of the old one-keyword-one-page playbook. One documented B2B example added roughly 30 new ranking keywords from a single-page rewrite that added eight sub-query H2 sections.

**The trap to avoid:** chasing individual fan-out queries as if they were keywords. The sub-queries are probabilistic and personalized — they differ between runs and users. As one analysis put it, optimize for the *topical themes* the fan-out reveals in aggregate, not the specific strings; chasing hyper-long-tail fragments one by one burns budget for nothing. Cover the cluster; don't stalk the fragments.

## Query fan-out analysis: a 10-minute method that needs no tool

People searching for a "query fan-out tool" usually want one thing: to see the sub-queries an AI generates for *their* topic. You can get most of the way there by hand, and the manual version teaches you more than a dashboard would.

1. **Start from the customer's real question**, not a keyword. "Best accountant for a small e-commerce business" — the way a person asks, not "small business accountant."
2. **Ask Google's AI Mode and expand the sources panel.** AI Mode shows the searches it ran to build the answer; those are literal fan-out sub-queries. Write them down verbatim.
3. **Ask ChatGPT with search on, then ask it:** "List the searches you ran to answer that." It will usually tell you — the list is the second engine's fan-out.
4. **Ask Perplexity the same question.** Its step-by-step "searching for…" lines are the third set.
5. **Cluster the three lists into themes.** You'll typically find 5–8 themes: a definition, a comparison, a price/cost angle, a "for my situation" angle, a how-to, a "what to avoid," and a local/availability angle. The themes recur between runs even when the exact strings don't — which is why themes, not strings, are the optimization target.
6. **Grade your page against each theme.** For every theme: is there a clearly-headed section on your page that answers it in its first two sentences? Missing themes are your to-do list; the most commonly missing are comparison and pricing.

That's the whole analysis. It takes ten minutes per topic, and it's what dedicated fan-out tools automate — usually by prompting a language model to *simulate* the sub-queries, which is a guess at what the engines do rather than an observation of it. Manual extraction from AI Mode's own "searches it ran" is closer to ground truth than any simulator.

**What about "grounding queries"?** Same idea, Google's vocabulary: when Gemini or AI Mode grounds an answer in search, the searches it issues are grounding queries — the fan-out set. If you see the term in Google's developer documentation, read it as "the sub-queries the model actually ran."

If you'd rather check the extractability side across your whole site at once — whether pages open with an answer, use question-shaped headings, and cite sources — that's what [SeeGeo's free audit](https://see-geo.com/) scores in about 20 seconds.

## How does this change what "success" looks like?

It moves measurement from *position* to *presence*. There is no position #1 inside a synthesized answer — there's only how often you appear across many answers to many related prompts. That means tracking citation rate and share-of-voice across engines, sampled repeatedly (AI answers vary run to run), instead of a daily rank number. Standard rank trackers don't capture fan-out exposure at all — which is precisely the measurement gap [SeeGeo](https://see-geo.com) is built for: the free audit scores whether your key pages are structured the way fan-out retrieval rewards — answer-first sections, liftable passages — and tracking measures how often AI answers actually name you, sampled repeatedly over time.

The one-sentence takeaway: your customer asks one question, the AI asks twenty on their behalf, and the business that answers the most of those twenty — clearly, near the top of the page, in liftable sections — is the one that gets named.

---

## Frequently asked questions

**What does query fan-out mean in simple terms?**
When you ask an AI search system one question, it secretly breaks it into many smaller related searches, runs them all at once, and builds its answer from whichever sources best answer each piece. Your content competes against those hidden sub-questions, not just the words the user typed.

**Is query fan-out the same as keyword variations?**
No. Traditional query expansion broadens a search with similar keywords; fan-out generates distinct, focused sub-questions targeting different facets and intents — including comparisons and implicit needs the user never typed. It's the difference between synonyms and follow-up questions.

**Does query fan-out replace traditional SEO?**
It builds on it. Backlinks, technical health, and page authority still heavily influence which sources get retrieved — roughly half of AI Overview citations come from top-10 ranking pages. Fan-out optimization is additive: strong SEO gets you into the retrieval pool; sub-query coverage gets you cited from it.

**How many sub-queries does an AI generate per question?**
It varies by engine and question complexity — from a handful to dozens. Google's AI Mode is observed to fan out most aggressively, ChatGPT moderately, Perplexity most narrowly. The exact set also changes between runs, which is why optimizing for aggregate themes beats chasing specific sub-query strings.

**How do I know if my page covers the fan-out for my topic?**
Collect the sub-queries (AI Mode's "searches it ran," ChatGPT extraction extensions, People Also Ask), then check whether your page has a clearly-headed section directly answering each theme within its first sentences. If comparative and pricing sub-questions are missing — the most commonly skipped types — start there. An automated extractability audit can score this across your whole site at once.

**Is it "query fan-out," "query fan out," or "query fanout"?**
All three mean the same thing and you'll see all three spellings; Google's own documentation hyphenates it ("fan-out"). Search engines treat the variants as equivalent, so the spelling you use on your own page doesn't matter — the coverage of the sub-questions does.

**Is there a query fan-out tool I should use?**
Several exist, and most work the same way: they prompt a language model to generate the sub-queries it *thinks* an engine would run. That's a simulation. The more reliable source is the engines themselves — Google's AI Mode shows the searches it ran, ChatGPT will list its searches when asked, and Perplexity displays them as it works. The 10-minute manual method above uses those directly; a tool is a convenience for doing it at scale, not a different source of truth.

**What are grounding queries, and how do they relate to fan-out?**
"Grounding queries" is Google's term for the searches Gemini or AI Mode issues when it grounds an answer in live search results. They are the fan-out sub-queries, seen from the developer-API side. If a page answers the themes those queries cover, it is a candidate to be cited in the grounded answer.

---

# We ran our own audit and scored a C. Here's the afternoon it took to get a B.
Source: https://see-geo.com/blog/we-audited-ourselves · Updated 2026-08-14

> SeeGeo audited itself: grade C, 78/100. What the engine caught, what an afternoon of fixes moved (entity clarity 66 → 92), and what code can't fix.

**We pointed our own audit engine at see-geo.com and it gave us a C — 78/100. One afternoon of fixes later, a re-run scored B, 82/100, with entity clarity jumping from 66 to 92. This is the full report, unedited: what the engine caught, what moved, and the part every audit tool glosses over — the findings no code change can fix. We're publishing it because a visibility tool that won't show you its own visibility isn't one you should trust.**

Every number below comes from the same free audit that runs on our homepage, plus our visibility scanner. Same engine, same scoring, no house discount.

## What did our own engine catch us doing?

The first run returned grade **C (78/100)** — SEO 85, GEO 70 — with eleven findings. The category breakdown told the story at a glance:

| Category | Score |
|---|---|
| Crawlability & access | 100 |
| Technical foundation | 85 |
| Structured data | 92 |
| Content extractability | **59** |
| Off-site presence & mentions | **55** |
| Entity clarity | **66** |

The plumbing was fine. The three weak categories were all about *meaning* — and two findings were genuinely embarrassing for a company that sells this exact fix:

**"Only machines can see what you do."** Our JSON-LD described SeeGeo precisely. Our visible homepage never did. A human — or any AI reading rendered text instead of structured data — had to work out what the product was. We'd made the classic mistake of telling the machines and forgetting the people.

**No About page.** The engine calls this "a trust gap in both Google's quality guidelines and AI systems' source weighing." It probed /about, /about-us, and /company. All dead. We sell entity clarity and had no entity page.

**Six French articles orphaned.** In the sitemap, linked from nowhere. Crawlers deprioritize pages nothing points to; ours literally couldn't be found by following links.

Plus: no visible dates anywhere ("freshness: 0/100"), no question-shaped headings on pages that answer questions, and our services page scored 56/100 on "can an AI quote this?"

## What did one afternoon of fixes buy?

We did exactly what our own reports tell customers to do, in report order:

- **Built the About page** (English and French) — answer-first opening, question-shaped headings, a visible date, and only checkable claims. We drafted a location line, couldn't verify it against anything, and deleted it. If it's not checkable, it doesn't go on a trust page.
- **Put the one-sentence definition in visible prose** on both landing pages: what SeeGeo is, in the first screen, for humans and rendered-text readers alike.
- **Added honest "Updated August 2026" lines** — from a manually-bumped constant, not an auto-stamp. An automatic date would claim freshness the content might not have, which is precisely the kind of trick the audit exists to catch.
- **De-orphaned the French blog** with a recent-articles section on the French landing page.
- **Rephrased headings as questions** where the page genuinely answers one.

Re-run, same engine, same rules:

| | Before | After |
|---|---|---|
| **Grade** | C (78) | **B (82)** |
| SEO | 85 | 89 |
| GEO | 70 | 75 |
| Entity clarity | 66 | **92** |
| Content extractability | 59 | 67 |
| Findings | 11 | 8 |

An afternoon. No redesign, no migration, no consultant. The single biggest mover — entity clarity, +26 — came from an About page and one visible sentence.

## What can't code fix?

Here's the section most tools won't write. Off-site presence scored 55 before the fixes and **55 after**, because no change to your own website makes other websites mention you.

The engine's top finding, before and after, was this: for our own category queries — things like "free tool to check if AI can see my website" — Gemini builds its answers from about ten specific pages, and **we appear on none of them**. It cites tools like searchscore.io, frase.io, and crawlercheck.com. We are, precisely, the problem we sell the fix for: reachable, readable, and absent from the answer.

The fix isn't code. It's getting named on the pages AI already trusts — directories, comparison posts, community threads. That's outreach work, it's slow, and our own engine hands us the exact target list. We'll report how it goes.

## The bigger lesson: being mentioned is not being read

While auditing ourselves, we ran a visibility scan on a brand at the other end of the spectrum — Linear, the issue tracker, a category leader. Twenty prompts, three runs each, live against Gemini:

- Mentioned in **59 of 60** answers (98%)
- Average position **1.4** when the answer was a ranked list
- Ahead of Jira on share of voice

Then we ran our citation probe, which records the web pages a search-grounded model actually retrieves while answering those same prompts. Across every answer that drew on live search, **linear.app appeared zero times**. What got read instead: comparison listicles (one project-management blog was read five separate times), YouTube, Reddit, and competitors' comparison pages.

Sit with that: the category leader is *in* nearly every AI answer, and the AI never once read the leader's own website to write them. The answers are assembled from third-party pages. Being the best product puts your *name* in the answer; being on the listicles is what puts your *pages* in the pipeline — and that's what you can lose, or win, without your product changing at all.

That asymmetry is the entire case for measuring this stuff instead of guessing.

## What happens next?

Our domain is days old, so search indexing is a waiting game — the honest accelerants are inbound links and time, not technical tricks. We'll keep working the target list our own audit produced, and we're building a public page that charts our own category visibility week by week, measured by our own scanner. When the line moves — either direction — you'll see it.

## Update — one week later: B 86

We re-ran the audit on August 16, a week after this article's C 78. The
score is now **B 86** — SEO 93, GEO 79 — and the findings list is down
from eleven to three. What moved it: entity clarity 66 → 92 (the About
page and visible definitions from the first pass), extractability 59 → 72
(answer-first landings now exist in all three languages), and structured
data at 92 with FAQ and HowTo markup across the blog.

Two honest footnotes. First, one "finding" from the earlier runs turned
out to be our engine's own false positive — it suggested FAQ markup on
the blog index, where the question-shaped headings are links to articles
rather than on-page answers. We fixed the engine, not the page: schema
that claims answers a page doesn't contain is exactly what we tell
customers not to publish. Second, everything still on the list is
off-site — the ten pages AI engines cite for our category where we don't
appear, and Wikipedia. (Wikidata: done — the item went up today and the
finding cleared within minutes.) Code can't fix those; outreach can. That's
the next chapter, and we'll publish it when the numbers move.

## Frequently asked questions

**Isn't publishing a C embarrassing?**
Less embarrassing than hiding it. The audit is deterministic — same site in, same score out — so the score was true whether we published it or not. A visibility company that curates its own visibility numbers is asking you to trust a mirror it won't look into.

**Why did the score only move 4 points if entity clarity jumped 26?**
Because the overall grade weighs six categories and the biggest remaining weakness — off-site presence — is immune to on-site fixes. That's honest scoring working as intended: an afternoon of edits shouldn't buy an A.

**Can I run the same audit on my site?**
Yes — it's the free audit on [our homepage](https://see-geo.com/#audit). Same engine, same scoring, usually under 30 seconds, first three audits free with no signup.

---

# Does AI visibility actually drive revenue? What the 2026 data shows
Source: https://see-geo.com/blog/ai-visibility-revenue-data-2026 · Updated 2026-08-13

> AI-referred visitors convert at roughly 4.4x organic on average. Here's the 2026 data on AI search traffic, revenue, and what it means for your business.

**Quick answer: yes — and the evidence is no longer anecdotal. Across independent 2026 studies, visitors who arrive from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush's cross-industry analysis), with individual studies reporting anywhere from 31% higher (ecommerce) to 23x higher (a B2B SaaS case from Ahrefs' own data). The volume is still small — about 1% of total website traffic across industries — but it's growing at a pace organic search has never seen, and it's the highest-converting traffic source most businesses aren't measuring at all.**

This article lays out the actual numbers, where they come from, what's solid versus what's inflated, and what it means for a small business deciding whether AI visibility is worth real budget.

## How much better does AI search traffic convert?

The headline finding is remarkably consistent across independent sources, even though the magnitudes vary:

| Study / Source | Finding | Scope |
|---|---|---|
| Adobe Digital Insights (Q1 2026) | AI-assistant visitors converted **42% better** than non-AI traffic in March 2026 — a full reversal from a year earlier, when the same channel converted 38% *worse* | Large-scale ecommerce panel |
| Semrush (2026) | AI-driven visitors convert at **4.4x** the rate of standard organic | Cross-industry |
| Seer Interactive | ChatGPT referrals converted at **15.9%** vs 1.76% for Google organic (~9x) | B2B-heavy client data |
| Ahrefs (internal) | 0.5% of visitors from AI search drove **12.1% of total signups** — a ~23x multiplier | Single company (their own site) |
| Visibility Labs (2025) | ChatGPT referrals converted at 1.81% vs 1.39% non-branded organic — **31% higher** | 94 ecommerce brands |
| Microsoft Clarity | Copilot referrals converted at **17x** the rate of direct traffic | 1,277 publisher domains |
| First Page Sage (May 2025–Feb 2026) | Largest lifts in hotels (+3.4 points, 3.6%→7.0%), higher education, legal, manufacturing | 160+ client companies |

Read that table honestly and two things are true at once. First, the *direction* is unanimous: every independent measurement finds AI-referred visitors converting meaningfully better than organic search visitors. Second, the *magnitude* is all over the map — from 1.3x to 23x — because industries, conversion definitions, and measurement windows differ wildly. The famous "23x" number is one company's internal data; the responsible planning figure is Semrush's cross-industry 4.4x, with ecommerce at the modest end and B2B SaaS at the dramatic end.

## Why do AI-referred visitors convert so much better?

Because the comparison shopping already happened inside the AI conversation. When someone asks ChatGPT "what's the best project management tool for a small agency" and then clicks through to a recommended brand, they've already seen that brand evaluated against alternatives in a synthesized answer. The click comes *after* the consideration phase, not before it. A traditional organic click, by contrast, is often the *start* of a research process that sprawls across ten tabs and three Reddit threads.

The behavioral data backs this up: SE Ranking found AI visitors spend 68% more time on websites than traditional search visitors. These aren't casual browsers — they arrive pre-qualified, having effectively delegated the shortlisting to the AI. Go Fish Digital, documenting their own agency's GEO program, described AI search as "operating as a sales agent, pre-qualifying users before they even reach our site."

## How big is this channel, really?

Small — and this is where a lot of coverage oversells. Conductor's cross-industry benchmark puts AI referral traffic at about **1.08% of total website traffic** on average in 2026. If someone tells you AI search has replaced Google, they're a year or three early.

But three facts make that 1% worth taking seriously:

**It's compounding fast.** Previsible's analysis of 6.77 million LLM-driven sessions across 166 GA4 properties — the largest study of its kind — found AI-referred sessions grew **9.9x in 19 months** (November 2024 to May 2026). SE Ranking documented ChatGPT's referral share jumping 36.7% in May 2026 alone, hitting an all-time high across the US, UK, and EU. Adobe tracked AI traffic to US retailers rising 393% in Q1 2026 year-over-year. No mature channel grows like this.

**The visible number understates the real number.** The Digital Bloom estimates that around **70% of AI-influenced traffic arrives without referrer data** — it shows up in your analytics as "direct" or as a branded Google search, because the common behavior is: get a recommendation from ChatGPT, then Google the brand name to find the site. Your analytics attribute that sale to branded search; ChatGPT actually made it. Post-purchase "how did you hear about us" surveys consistently surface AI influence that dashboards miss.

**It's concentrated, which simplifies strategy.** Previsible found ChatGPT commands **92.4% of trackable LLM referral traffic**; Conductor puts it at 87.4%. Gemini is the quiet #2 and growing; Perplexity has fallen 61% from its March 2025 peak. For a small business, this means the game is less fragmented than the tool-vendor marketing suggests: win ChatGPT visibility first, cover Gemini and Google's AI surfaces second.

## Does AI visibility build brand awareness too, or just traffic?

Both — and the awareness effect may matter more, precisely because of that attribution gap. SE Ranking's analysis of the May 2026 referral surge found something telling: AI engines disproportionately send users to **homepages** rather than deep pages. People aren't arriving on a blog post; they're arriving on the brand itself, often hearing of it for the first time inside an AI answer. That's a brand-introduction channel, not just a traffic channel.

The measurable proxy is branded search volume: businesses that improve AI visibility consistently report branded Google searches rising in parallel — people see the recommendation, then search the name. If you track one "awareness" metric alongside referral traffic, track that.

## What should a small business actually do with these numbers?

Three practical moves, in order:

**1. Start measuring before you start optimizing.** Most GA4 setups bucket ChatGPT and Perplexity referrals under "Other" or "Direct." Segment AI referral sources explicitly, watch conversion rate by source rather than raw volume, and add a "how did you find us" field to signup or checkout. Only 14% of marketers track AI search performance at all — measuring it is, briefly, a competitive advantage in itself.

**2. Weight the channel by value, not volume.** At ~1% of traffic converting at ~4.4x, AI referrals are plausibly worth ~4–5% of your conversions today — and growing 10x every year and a half. That justifies real but proportionate effort: making sure AI crawlers can access your site, structuring content to be citable, and building presence in the sources AI engines cite. It does not justify abandoning SEO — especially since AI engines draw heavily on content that ranks well anyway.

**3. Judge success by trend, not snapshots.** AI answers are non-deterministic; your visibility is a percentage across many prompts over time, not a rank. Set a baseline, re-measure monthly, and correlate visibility changes with branded search and AI-referred conversions.

## The caveats we'd want you to know

Much of the loudest data in this space is published by tools and agencies selling GEO services (including, yes, us — SeeGeo sells AI-visibility software, and you should weigh this article accordingly). That's why this piece leans on the most independent sources available: Adobe's panel data, Semrush's cross-industry dataset, Microsoft Clarity's publisher study, and Previsible's multi-property GA4 analysis. Where a number comes from a single company's internal data (Ahrefs' 23x) or a vendor's client base, we've said so. The honest summary: the conversion premium is real and replicated; the exact multiple for *your* business is unknowable until you measure it.

That measurement is what SeeGeo automates — the [free audit](https://see-geo.com/) shows whether AI systems can see you at all, and the tracking plans measure whether they mention you, how you're described, and whether that's trending up. But whether you use our tool or a spreadsheet: the data above says this channel is no longer optional to at least *watch*.

---

## Frequently asked questions

**What percentage of website traffic comes from AI in 2026?**
About 1.08% on average across industries, per Conductor's benchmark — but growing roughly 10x every 19 months (Previsible), and understated by attribution: an estimated 70% of AI-influenced visits arrive without referrer data, showing up as direct or branded-search traffic instead.

**Which AI platform sends the most referral traffic?**
ChatGPT, by a wide margin — 87–92% of trackable AI referral traffic depending on the study. Gemini is the growing second; Perplexity's share has declined significantly since early 2025.

**Is AI search traffic really worth optimizing for at 1% of visits?**
For most businesses, yes — because that 1% converts at roughly 4.4x organic on average (with studies ranging from 1.3x to 23x by industry), spends 68% more time on site, and the channel is growing faster than any comparable one. The proportionate response is measurement plus foundational optimization, not a full budget pivot.

**How do I track ChatGPT referral traffic in Google Analytics?**
Create a segment or channel group matching referral sources containing chatgpt.com, chat.openai.com, perplexity.ai, gemini.google.com, and copilot.microsoft.com. Then compare conversion rate — not just sessions — against your organic baseline. Add a "how did you hear about us" survey to capture the AI influence that arrives without referrer data.

**Does being mentioned by AI increase brand awareness even without clicks?**
Yes — and it's partly measurable. AI recommendations frequently lead to branded Google searches rather than direct clicks, so rising branded search volume alongside improving AI visibility is the standard proxy for the awareness effect. AI engines also disproportionately link to homepages, introducing the brand rather than a single page.

---

# 7 real examples of businesses that turned AI visibility into revenue
Source: https://see-geo.com/blog/geo-case-studies-real-examples · Updated 2026-08-13

> A restaurant with 520 bookings, a SaaS with 32% of leads from ChatGPT, a B2B firm with 4,900% revenue growth. Real GEO case studies — and what each did.

**Quick answer: businesses across sizes and industries are now documenting real revenue from AI visibility — a building-products company grew Google AI Overviews mentions 540% alongside a 67% organic traffic lift; prop-tech SaaS Smart Rent reports 32% of its sales-qualified leads originating from ChatGPT citations; a local restaurant attributes 520 booking-form submissions to AI-driven discovery; and one B2B tech client documented 4,900% revenue growth from LLM-referred sources over 14 months.** One honest caveat before the examples: nearly all published GEO case studies come from the agencies and tools that ran them, so treat the exact figures as self-reported. What's more reliable — and more useful — is the *pattern* in what the winners actually did, which repeats across every case below.

## Case 1: The building-products company that became Google's AI answer

**LS Building Products** (US building materials) is one of the most-cited GEO results in circulation, with figures reported across multiple industry write-ups: a **540% increase in Google AI Overviews mentions**, a **67% increase in organic traffic**, and a **400% rise in traffic value** over roughly six months.

**What they actually did:** translated deep product expertise into plainly-written, well-structured explanatory content — the kind an AI can lift and cite. Question-shaped pages, direct answers, accessible language instead of industry jargon.

**The lesson for a small business:** you don't need new expertise — you need your existing expertise restructured into liftable answers. The knowledge was already in the company; the 540% came from repackaging it.

## Case 2: The SaaS getting a third of its qualified leads from ChatGPT

**Smart Rent**, a prop-tech SaaS serving property managers, recognized that its enterprise buyers had stopped typing keyword fragments and started asking AI platforms detailed questions. Reported results: **32% of sales-qualified leads originating from ChatGPT citations**, alongside a 200% rise in AI-search-driven traffic and a 32% lead increase overall.

**What they actually did:** optimized for the *questions* buyers ask AI ("how do property managers handle smart-home access for tenants?") rather than the keywords they used to type, and structured content so AI engines could reference it when answering those exact questions.

**The lesson:** in B2B especially, AI search functions as a pre-qualifier — Smart Rent's AI-referred leads arrived already educated. If your buyers research before contacting you, the research is increasingly happening inside an AI chat.

## Case 3: The agency that measured a 25x conversion premium on itself

**Go Fish Digital**, an SEO agency, ran GEO on its own business and published the numbers: roughly **3x lead growth**, with AI-referred leads converting at **25x the rate** of traditional search leads. Their framing is the most quotable summary of the channel to date — AI search acting as a sales agent that pre-qualifies users before they ever reach the site.

**What they actually did — and the crucial admission:** they note openly that they started with an unusual advantage: years of existing citations, reviews, Reddit presence, and third-party recognition, meaning LLMs already "knew" them. Their GEO work built on an existing mention footprint.

**The lesson:** external credibility signals — reviews, mentions, community presence — are the raw material LLMs work from. A business with no third-party footprint should expect to build that foundation *first*; a business that has one can see results much faster.

## Case 4: The B2B client with 4,900% revenue growth from LLM traffic

The most dramatic documented result comes from **The Optimist**, a B2B content agency, for a technology client: over a 14-month engagement, **4,900% revenue growth and 2,622% traffic growth from LLM-referred sources**. (Percentages that large usually mean a small starting base — which is exactly the situation most businesses are in with AI traffic today.)

**What they actually did:** built the entire strategy around **original first-party research** — proprietary studies and datasets that LLMs cite as primary sources, rather than repurposed industry commentary.

**The lesson:** AI engines need sources for claims, and original data makes you *the* source instead of one paraphrase among many. For a small business this scales down honestly: survey your customers, publish the numbers, become the citation for one specific claim in your niche.

## Case 5: The restaurant and the venue — proof this isn't just for SaaS

Two local-business examples from published AEO case-study collections: **The Albert**, a restaurant, attributes **520 booking-form submissions** to AI-driven discovery; **444Social**, an events venue, reports reaching **100% occupancy** with AI visibility as a contributing channel — both on local-business budgets in the $1–3K/month range rather than enterprise retainers.

**What they actually did:** the local playbook — consistent name/address/hours everywhere, review presence, clearly structured menus and event pages AI can read, and content answering the questions people actually ask assistants ("private dining room for 20 in [city]").

**The lesson:** local intent has moved to AI faster than almost anyone predicted — BrightLocal's 2026 consumer survey found AI use for local business recommendations jumped from 6% to 45% in a single year. For local businesses, the fixes are cheap, and the channel is suddenly a top-three discovery path.

## Case 6: The beauty brand that 3.3x'd its AI mentions in 60 days

A global haircare brand (documented by the GEO platform OptimizeGEO, so vendor-reported) grew total AI mentions from **86 to 282+ in 60 days** — moving from absent-or-inconsistent to consistently present across its core problem categories (hair fall, frizz, damage repair).

**What they actually did:** two notable tactics beyond the standard playbook — **YouTube transcript optimization** (Gemini pulls directly from video transcripts, making existing video content a citation source) and a **quarterly freshness cadence** for refreshing key content, plus building presence on the platforms where AI models source real-world opinions.

**The lesson:** your citable surface is bigger than your website. Video transcripts, community threads, and review platforms all feed AI answers — and 60 days is enough to move mention counts when the gaps are structural rather than reputational.

## Case 7: The mid-market software company that fixed the invisible-despite-ranking problem

A project-management software provider (documented in a 2026 GEO case collection) had strong traditional Google rankings but near-zero AI presence — the exact "ranking but invisible" gap. Over six months: **340% increase in AI search mentions and 67% more qualified demo requests**.

**What they actually did:** restructured existing high-ranking content into answer-first format, added comparison content matching how buyers phrase questions to AI, and implemented structured data across the content library — no new topics, just new shape.

**The lesson:** ranking well on Google and being cited by AI are correlated but separate outcomes. If you already rank, you've done the hard part; the AI layer is largely a formatting and structure project on top of assets you own. [Checking whether you have that gap](https://see-geo.com/blog/is-your-website-visible-to-chatgpt) takes about five minutes.

## What do all seven have in common?

Strip the numbers away and the same five moves appear in every single case:

1. **Answer-shaped content** — direct answers to real questions, in plain language, structured for extraction.
2. **Third-party presence** — reviews, mentions, community threads, and citations that teach LLMs the brand exists and can be trusted (the factor Go Fish credits for their head start).
3. **Original, citable substance** — proprietary data, genuine expertise, or first-hand specificity that makes the brand a source, not a paraphrase.
4. **Technical accessibility** — AI crawlers allowed in, structured data in place, content readable without JavaScript.
5. **Measurement over time** — every documented winner tracked mentions/citations as a trend, not a one-time check.

That list is, not coincidentally, the exact sequence a good SEO+GEO audit checks — it's the checklist [SeeGeo's free audit](https://see-geo.com/) runs automatically. But whether you use a tool or do it by hand: the case studies above are unanimous that the work is knowable, repeatable, and — [per the conversion data](https://see-geo.com/blog/ai-visibility-revenue-data-2026) — increasingly the highest-value traffic most businesses aren't competing for yet.

---

## Frequently asked questions

**Are GEO case study results trustworthy?**
Directionally yes, precisely no. Nearly all published GEO case studies are self-reported by the agencies or tools involved, and dramatic percentages often reflect small starting bases. The consistent cross-case pattern (answer-shaped content + third-party presence + original substance + technical access) is more reliable than any individual number. Independent conversion data — Adobe's 42% premium, Semrush's 4.4x — corroborates that the underlying channel is real.

**How long did these results take?**
The documented range: 60 days (beauty brand mention growth) to 14 months (the 4,900% revenue case), with 6 months as the most common timeline for meaningful mention and lead growth. Access and formatting fixes show effects fastest; mention-footprint building compounds over quarters.

**Can a small local business really benefit from GEO?**
Yes — arguably fastest of anyone. AI use for local recommendations grew from 6% to 45% of consumers in one year (BrightLocal, 2026), the local playbook (consistent listings, reviews, structured pages) is inexpensive, and documented local cases like The Albert's 520 bookings ran on $1–3K/month budgets, not enterprise retainers.

**What's the single highest-leverage tactic across these cases?**
For businesses with an existing web presence: restructuring current content into answer-first, extractable format (Cases 1 and 7 got their entire results this way). For businesses AI doesn't know yet: building third-party mentions and reviews first, since that footprint is what LLMs learn brands from.

**How do I know if my business is even eligible to be cited?**
Check whether AI crawlers can access your site (robots.txt, CDN settings, JavaScript dependence) and whether AI platforms describe your business accurately when asked. Both take minutes to check manually — or run a free automated audit that covers eligibility and content structure together.

---

# Why does ChatGPT recommend my competitor instead of me? (and how to fix it)
Source: https://see-geo.com/blog/chatgpt-recommends-competitor-not-me · Updated 2026-08-11

> If ChatGPT or AI Overviews keep naming your competitor and not you, one of six fixable causes is usually why. How to diagnose and close the gap.

**Quick answer: when an AI assistant consistently recommends your competitor and not you, it's almost never random — it usually traces to one of six fixable causes: your site blocks AI crawlers, your content doesn't answer questions in a liftable way, your competitor appears in the third-party sources AI engines cite, your business is described inconsistently across the web, your competitor simply has more of a mention footprint, or the AI can't confidently categorize what you do.** All six are diagnosable in under an hour, and the fixes range from a one-line file edit to a few months of steady mention-building.

This one stings more than a Google ranking gap, and for good reason: a Google page shows ten results, so losing still means being seen. An AI answer often names one to three businesses — or just one. When that one is your competitor, your customer may never learn you exist. With a 2026 study cited by Search Engine Land finding that 37% of consumers now begin their research in AI tools rather than search engines, that's not a hypothetical channel anymore.

Here's how to figure out which cause is yours, in the order you should check them.

## Reason 1: Your site blocks AI crawlers — and theirs doesn't

The most brutal and most common cause. AI companies read the web with their own crawlers — [GPTBot](https://see-geo.com/bots/gptbot) (OpenAI), [ClaudeBot](https://see-geo.com/bots/claudebot) (Anthropic), [PerplexityBot](https://see-geo.com/bots/perplexitybot), [Google-Extended](https://see-geo.com/bots/google-extended) — and your site can block them in its robots.txt file, or a security layer like Cloudflare can turn them away by default without your robots.txt showing anything. Meanwhile Google sees you fine, so your rankings look healthy and nothing seems wrong.

**Check it:** visit `yourdomain.com/robots.txt` and look for `Disallow: /` rules under any of those bot names. Then do the same for your competitor's robots.txt (it's public — just type their domain + /robots.txt). If you block and they don't, you've likely found your answer. Our [5-minute visibility check](https://see-geo.com/blog/is-your-website-visible-to-chatgpt) walks through this step by step.

**Fix:** remove the blocking rules (or change the CDN setting). This is a minutes-long fix with the highest leverage of anything in this article, because every other factor is irrelevant while the door is locked.

## Reason 2: Their content is liftable — yours is browsable

AI engines don't rank your page; they *extract from* it. A page that opens with three paragraphs of brand story before getting to the point gives an AI nothing to lift. A page that opens with "X costs between $A and $B, depending on Y" hands the AI its answer, pre-written and attributable.

**Check it:** open your most important page and your competitor's equivalent side by side. Ask: within the first two paragraphs, does each page directly answer the question a customer would ask? Are headings phrased as questions? Are there concrete numbers, tables, dates? Princeton's research on generative engine optimization found that adding statistics and citing sources improved visibility in AI answers by roughly 30–40% — if their page is stat-rich and yours is adjective-rich, that's a measurable gap, not a stylistic one.

**Fix:** restructure key pages answer-first: direct answer up top, question-shaped H2s, real numbers with sources, comparison tables, visible dates and a named author. This is rewriting, not rebuilding — usually days of work, not months. Our [GEO writing checklist](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt) is the exact nine-point version we apply.

## Reason 3: They're present in the sources the AI already cites — you're not

This is the least obvious cause and often the decisive one. When ChatGPT or Perplexity answers "best [your category] in [your city]," it leans on third-party pages: "best of" listicles, directories, review platforms, Reddit threads. If your competitor appears in five of the pages the AI habitually cites and you appear in none, the AI is faithfully reflecting its sources. The competition happened upstream, on pages neither of you owns.

**Check it:** ask ChatGPT, Perplexity, and Google's AI Overviews your customers' top three questions. Note every source cited. Open each one and search it for your name and your competitor's name. Tally it up — this is your citation-source gap, and it's usually the single most explanatory number in the whole diagnosis.

**Fix:** work the list. Getting added to an existing "best of" article, claiming and enriching directory profiles, earning presence on the review platforms in your category, participating credibly where your market discusses options — these are outreach and reputation tasks, not technical ones. They're slower than a robots.txt edit, but each placement compounds because it feeds every future AI answer drawing on that source.

## Reason 4: The web describes your competitor consistently — and you inconsistently

AI systems assemble an understanding of your business from everything written about it. If one directory calls you a "marketing agency," another an "SEO consultancy," your old Facebook page says "web design studio," and your homepage says "growth partner," the model's picture of you is blurry. Blurry entities don't get confidently recommended; models recommend what they can cleanly categorize.

**Check it:** ask ChatGPT "What is [your business name]?" and see whether the answer is accurate, outdated, or confused. Then Google your own business name and read your top ten mentions the way a stranger would — do they tell one story?

**Fix:** align everything. One consistent name, one clear category description, one accurate location, on your site's homepage and About page and across every directory and profile you control. For local businesses this extends to NAP consistency (name, address, phone) everywhere it appears. Unglamorous work; disproportionate payoff.

## Reason 5: They simply have a bigger mention footprint

Sometimes there's no trick to find — your competitor has been written about more, reviewed more, discussed more. Classic SEO counted backlinks; AI systems absorb *mentions*, linked or not, so a competitor with years of press hits, podcast appearances, and community presence has a lead that no technical fix erases.

**Check it:** compare mention volume honestly: search both brand names, check Reddit and industry forums, count reviews on the platforms your category uses.

**Fix:** you close this gap the slow way — steadily generating things worth mentioning (original data, useful free tools, genuinely helpful content) and showing up where your market talks. The encouraging part: AI answers have a recency lean, and a smaller brand with sharper, more extractable content and cleaner presence in cited sources regularly outperforms a bigger but lazier one. Footprint matters; it isn't destiny.

## Reason 6: The AI can't tell what you actually do

Related to Reason 4 but worth its own check because it's so common with SMB websites: homepages that lead with slogans ("We turn ambition into momentum") and never plainly state the category. A human squints and figures it out. A model that can't confidently classify you won't risk recommending you for anything specific.

**Check it:** show your homepage's first screen to someone who's never heard of you and give them five seconds. Can they say what you sell and who it's for? If not, neither can a crawler.

**Fix:** add one plain sentence high on the homepage: "[Name] is a [category] in [location] that [what you do] for [whom]." Keep the slogan; just don't make it do the categorization job alone.

## How do I know which reason is mine?

Run the checks in the order above — they're sequenced from fastest-to-check and highest-leverage downward. In practice: access issues (Reason 1) are pass/fail and explain the worst cases; content extractability (2) and citation-source gaps (3) explain most of the rest; entity problems (4 and 6) are the silent multiplier on everything; raw footprint (5) is the residual you grind down over quarters.

Or shortcut the hour of manual checking: [SeeGeo's free audit](https://see-geo.com/) runs the access, extractability, structured-data, and entity checks automatically in about 30 seconds, and the paid plan tracks the citation-source and share-of-voice comparison against competitors continuously — so "why them and not us" becomes a dashboard number that moves, instead of a mystery you re-investigate every quarter.

---

## Frequently asked questions

**Why doesn't ChatGPT know my business exists?**
Usually one of three reasons: your site blocks AI crawlers (check your robots.txt and CDN settings), your business has very little third-party web presence for the model to have learned from, or your site's content doesn't clearly state what the business is. All three are fixable; the first takes minutes.

**Can I pay to be recommended by ChatGPT or Perplexity?**
No. As of 2026 there is no placement you can buy inside the organic answers of the major AI assistants. Visibility comes from being accessible, extractable, and well-represented in the sources those systems draw on — which is why the work described in this article exists.

**How long does it take to catch up to a competitor in AI answers?**
Access and content fixes can show effects in retrieval-based engines like Perplexity within days to weeks. Presence-building — getting into cited sources, growing mentions, cleaning up entity consistency — typically shows movement over one to several months. Anyone quoting exact timelines is guessing; AI answers are non-deterministic and models update constantly, which is why measuring on a schedule beats checking once.

**My competitor is bigger. Is this fight winnable?**
Often yes, because AI answers reward extractability and clarity, not just size. A smaller brand with answer-first content, clean structured data, consistent descriptions, and presence in a handful of key cited sources frequently out-appears a larger competitor coasting on legacy authority. Size sets the difficulty, not the outcome.

**Does being recommended by AI actually drive customers?**
Increasingly. With around 37% of consumers starting research in AI tools and click-through rates on AI-summarized queries down as much as 61% since mid-2024, the recommendation inside the answer is capturing intent that used to arrive as website clicks. The channel is young enough that attribution is imperfect — but absence from it is a compounding disadvantage.

---

# How to write content that ChatGPT actually cites: the GEO writing checklist
Source: https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt · Updated 2026-08-11

> AI engines cite content they can lift, verify, and attribute. A 9-point checklist for structuring pages so ChatGPT and AI Overviews quote you.

**Quick answer: AI engines cite content they can lift cleanly, verify easily, and attribute confidently. In practice that means pages that open with a direct answer, use question-shaped headings, include concrete statistics with named sources, package comparisons in tables, show a date and a named author, and carry matching schema markup.** The Princeton research that founded generative engine optimization put numbers on this: adding statistics and citing sources improved visibility in AI answers by roughly 30–40% — making these the highest-leverage edits you can make to existing content.

This is the working checklist we apply to every page, including this one. Nine items, ordered by impact, each with the *why* and a before/after so you can apply it in an afternoon.

## How is writing for AI different from writing for Google?

One sentence of theory before the checklist: Google ranks whole pages and sends a human to browse them; AI engines extract *passages* and assemble them into an answer. Google's unit of competition is the page. The AI's unit of competition is the paragraph. So GEO writing is the craft of making individual passages self-contained, factual, and quotable — while keeping the page genuinely good for the humans who arrive. Nothing below trades one for the other; extraction-friendly writing is, mostly, just unusually clear writing.

## The 9-point GEO writing checklist

### 1. Open with the answer, not the wind-up

The first two paragraphs should directly answer the question the page targets — the complete short version, not a teaser. AI systems disproportionately draw from the top of documents, and an opening that actually answers gets lifted; an opening that promises to answer later gets skipped.

*Before:* "Choosing the right espresso machine can feel overwhelming. With so many options on the market, where do you even begin? In this guide, we'll explore everything you need to know…"

*After:* "For most home users, the best entry-level espresso machine in 2026 is a single-boiler semi-automatic in the $300–500 range — enough for café-quality shots without commercial complexity. Here's how to choose within that range, and when it's worth spending more."

### 2. Phrase headings as the questions people actually ask

Your H2s should read like queries: "How much does X cost?" "Is X better than Y for Z?" This isn't cosmetic — it tells the AI exactly which question the following passage answers, making the match between a user's question and your content trivially easy. It also mirrors how people phrase things to assistants, which is more conversational and specific than the two-word phrases they type into Google.

### 3. Add real statistics — and name where they came from

The single most evidence-backed item on this list. The Princeton GEO study found statistics and source citations were the top-performing content additions, each improving AI visibility on the order of 30–40%. The mechanism is intuitive: AI systems prefer repeating claims that arrive pre-verified. "Many customers now research with AI" is an opinion; "37% of consumers begin searches with AI tools, per a 2026 study cited by Search Engine Land" is a citable fact with a chain of custody.

Two rules: only real numbers (an invented statistic that spreads is a reputational time bomb), and name the source in the sentence itself, not just a hyperlink — the attribution needs to survive being lifted out of your page.

### 4. Write standalone passages of 2–4 sentences

Audit your key sections with one test: if this paragraph were quoted alone, with no surrounding context, would it still be true, clear, and complete? Passages stuffed with "as mentioned above," "this approach," and dangling pronouns die in extraction. Self-contained passages travel.

*Before:* "This makes it a great option for the use case we discussed, though it has the drawback mentioned earlier."

*After:* "A single-boiler machine suits home users who make one or two drinks at a time; its main drawback is waiting 30–60 seconds between pulling a shot and steaming milk."

### 5. Put comparisons in real tables

Anything comparative — options, prices, pros and cons, us-vs-them — belongs in an actual HTML table, not prose and not an image. Tables are the most cleanly extractable structure on the web: both AI engines and Google's rich results can read, reproduce, and attribute them. An image of a table, by contrast, is invisible to most extraction.

### 6. Show a visible date — and keep it honest

AI engines lean toward fresh sources, and a visible "Updated: [date]" is the signal they read. Add published and updated dates to your templates, and actually refresh cornerstone pages on a schedule (quarterly is a reasonable default in fast-moving categories). The honesty clause: bumping dates without changing content is detectable and erodes exactly the trust you're trying to build.

### 7. Attach a named author with a one-line credential

"By the Team" is an anti-signal. A named human with a stated credential ("15 years running local restaurant kitchens") feeds the E-E-A-T trust cluster that both Google and AI systems weigh when deciding which sources are safe to repeat. Add an author box to your templates; link it to a real bio page.

### 8. Mirror the structure in schema markup

Schema (structured data) is machine-readable labeling that confirms what your page contains. For GEO writing, three types matter most: Article (basic identity, author, dates), FAQPage (marks your question-answer pairs), and HowTo (marks step sequences — like this checklist). Schema doesn't rescue weak content, but it removes ambiguity from strong content, and ambiguity is what keeps borderline passages from being used. If your pages already follow items 1–7, the schema almost writes itself, because the content is already shaped like the markup.

### 9. End with an FAQ section that earns its place

A closing FAQ block of 4–6 real questions — the ones people actually ask next — does triple duty: it captures long-tail question searches, it maps one-to-one onto FAQPage schema, and each answer is a purpose-built liftable passage. The discipline: answer each in 2–4 self-contained sentences, and don't pad with filler questions nobody asks. (You'll find ours below, practicing the format.)

## What should I fix first on existing content?

Don't rewrite the whole site. Ruthless order of operations:

1. **Pick your five most important pages** — the ones answering the questions that bring customers.
2. **Apply items 1 and 3 first** (answer-first opening, statistics with sources) — highest measured impact per hour of work.
3. **Then items 2, 5, and 9** (question headings, tables, FAQ) — structural, template-level changes.
4. **Then 6, 7, 8** (dates, authors, schema) — mostly one-time template fixes that pay across every page.

An afternoon per page is a realistic budget for the first pass. And a necessary caveat this industry doesn't say enough: GEO is roughly two years old, resting on one landmark study and fast-moving practitioner consensus. These practices measurably help today; treat them as the current best playbook, not eternal law, and re-test as models change.

## How do I know if it's working?

You measure it the only way AI visibility can be measured: by asking. Build a panel of the 10–20 questions your customers actually ask, run them across ChatGPT, Perplexity, Gemini, and Google AI Overviews, and record whether you're mentioned, cited with a link, or actively recommended — then repeat monthly, because AI answers vary between runs and single checks are snapshots, not scores.

That loop is tedious by hand, which is precisely why we built it into [SeeGeo](https://see-geo.com/): the free audit scores your existing pages against this checklist automatically, and the paid plan runs your question panel on schedule so you watch citations trend instead of guessing. Either way — tool or by hand — measure. Writing without measuring is how this discipline stays folklore.

If the checklist isn't moving anything, the problem may be upstream: see [why ChatGPT recommends your competitor](https://see-geo.com/blog/chatgpt-recommends-competitor-not-me), where the cause is often access or citation-source presence rather than the writing itself.

---

## Frequently asked questions

**How do I get my website cited by ChatGPT?**
Make sure AI crawlers can access your site, then publish answer-first content: direct answers in the opening paragraphs, question-shaped headings, real statistics with named sources, comparison tables, visible dates, named authors, and matching schema markup. Then build presence in the third-party sources AI engines already cite for your topic, and measure monthly.

**Does schema markup help with AI citations?**
It helps by removing ambiguity: schema confirms in machine-readable form what your page contains, which supports extraction and attribution. It doesn't compensate for content that isn't answer-shaped — think of it as the label, not the product.

**How long should GEO-optimized content be?**
As long as the answer requires and no longer. Depth helps when it adds facts, examples, and covered sub-questions; padding hurts because it buries the liftable passages. A thorough 1,500-word page with ten self-contained passages beats a 4,000-word page with three.

**Should I write different content for AI engines and for Google?**
No — one page, written answer-first, serves both. The overlap between what AI engines extract and what Google's helpful-content systems reward is large, and maintaining parallel versions of your content is wasted effort with added risk of inconsistency.

**Can AI-generated content get cited by AI engines?**
There's no explicit penalty for how content was produced; the systems judge the output. In practice, generic AI-generated content fails this checklist on its own — no original statistics, no real author, no first-hand specifics — which is what keeps it from being cited. Content drafted with AI but grounded in your real data, expertise, and sources can absolutely perform.

**What's the fastest single change I can make today?**
Rewrite the opening of your most important page so its first paragraph fully answers the question the page targets, and add one real statistic with its source named in the sentence. Those two edits carry the strongest evidence behind them and take under an hour.

---

# Is your website invisible to ChatGPT? Here's how to check in 5 minutes
Source: https://see-geo.com/blog/is-your-website-visible-to-chatgpt · Updated 2026-08-11

> Many sites block AI crawlers without knowing it. A free 5-minute check to see whether ChatGPT, Perplexity, and Gemini can actually see your business.

**Quick answer: your website can be completely invisible to ChatGPT, Perplexity, and Gemini even if it ranks fine on Google — and it takes about five minutes to find out.** The most common causes are a robots.txt file that blocks AI crawlers (often set by default, not by choice), a security service like Cloudflare silently turning AI bots away, or a site built so heavily on JavaScript that AI crawlers see a blank page. This guide walks you through checking all of it yourself, step by step, with no tools required.

That matters more than it used to. A 2026 study cited by Search Engine Land found that 37% of consumers now begin searches with AI tools rather than traditional search engines, and organic click-through rates for queries that trigger AI summaries have fallen by as much as 61% since mid-2024. If a meaningful slice of your customers is asking ChatGPT "what's the best [your category] near me" — and your site is blocked — you're not losing to competitors. You're not even in the room.

## Why would my website be invisible to AI in the first place?

Because AI visibility and Google visibility are controlled by different switches, and most websites only ever flipped the Google one.

Search engines and AI companies each send their own crawler — a bot that reads your site. Google sends Googlebot. OpenAI sends [GPTBot](https://see-geo.com/bots/gptbot). Anthropic sends [ClaudeBot](https://see-geo.com/bots/claudebot). Perplexity sends [PerplexityBot](https://see-geo.com/bots/perplexitybot). Google's AI training crawler is separate from its search crawler and is called [Google-Extended](https://see-geo.com/bots/google-extended). Your website's robots.txt file (a plain text file that tells bots what they may read) can allow some and block others — and many sites block the AI ones without the owner ever deciding to.

There are three common ways this happens:

1. **A default you never chose.** Cloudflare, which sits in front of a huge share of the web, moved to blocking AI crawlers by default. If your site is behind Cloudflare and nobody actively changed that setting, AI bots may be getting turned away at the door — with your robots.txt looking perfectly innocent.
2. **A well-meaning past decision.** During the 2023–2024 wave of concern about AI companies training on website content, many site owners (or their agencies, or a WordPress plugin) added blanket AI-blocking rules. Reasonable at the time — but the trade has changed: blocking AI crawlers now also means being absent from AI answers your customers actually read.
3. **A site AI can't read.** Some AI crawlers do not run JavaScript. If your site is a JavaScript-heavy single-page app that only shows content after scripts run, an AI crawler can fetch your page and receive, functionally, nothing.

None of these show up in Google Analytics. Traffic looks normal. Rankings look normal. The invisibility is happening in a channel you're not measuring.

## How do I check if ChatGPT can see my website? (the 5-minute check)

You need a browser and nothing else. Do these four steps in order.

### Step 1: Read your robots.txt (2 minutes)

Type your domain followed by `/robots.txt` into your browser — for example, `yourbusiness.com/robots.txt`. You'll see a plain text file. Look for lines mentioning any of these names:

- `GPTBot` (OpenAI / ChatGPT)
- `ClaudeBot` (Anthropic / Claude)
- `PerplexityBot` (Perplexity)
- `Google-Extended` (Google's AI training)
- `CCBot` (Common Crawl, a dataset many AI systems learn from)

The pattern that matters looks like this:

```
User-agent: GPTBot
Disallow: /
```

That pair of lines means "GPTBot may read nothing." If you see it — or the same pattern for any bot above — that crawler is blocked from your entire site. If none of these names appear at all, the file isn't blocking them (though Step 3 can still be).

### Step 2: Ask the AI about your own business (1 minute)

Open ChatGPT (and Perplexity, if you have two spare minutes) and ask the questions a customer would actually ask:

- "What is [your business name]?"
- "Best [your category] in [your city]"
- "Alternatives to [your biggest competitor]"

You're checking three things: Do you appear at all? Is what it says about you *accurate*? And who appears instead of you? Write down which sources the AI cites — those pages are where AI already looks for answers in your category, and being mentioned on them is one of the fastest visibility wins available.

One caveat: AI answers vary between runs. Not appearing once doesn't prove you're invisible; appearing once doesn't prove you're reliably visible. Treat this as a smoke test, not a measurement.

### Step 3: Check whether a security layer is blocking bots (1 minute)

If you know your site uses Cloudflare or a similar service (your web person will know; a quick way to guess is that your DNS is managed there), log into that dashboard and find the bot-management or "AI crawlers" settings. Look for whether AI bots are being blocked at the network level. This is the sneaky one: it overrides whatever your robots.txt politely says, because blocked bots never get far enough to read it.

### Step 4: View your site the way a no-JavaScript crawler does (1 minute)

In Chrome: open your homepage, press F12, open the command menu (Ctrl/Cmd+Shift+P), type "Disable JavaScript," select it, and reload the page. Alternatively, use any online "fetch as bot" tool.

Now look at the page. Is your actual content there — your service descriptions, your products, your location — or is it a blank shell with a loading spinner? If the content is gone, AI crawlers that skip JavaScript are seeing that blank version. Your beautiful site is, to them, an empty room.

## What do my results mean?

- **All four steps look clean:** AI systems can at least *access* you. Whether they *choose* you is the next battle — that's about how your content is structured, whether it answers questions directly, and how often your brand is mentioned around the web. (That's [a different article](https://see-geo.com/blog/seo-vs-geo-difference).)
- **Step 1 or 3 showed blocking:** you have a critical, fixable problem. Unblocking is usually a one-line robots.txt edit or one dashboard toggle, and it's the single highest-leverage fix in AI visibility because everything else depends on access.
- **Step 4 came back blank:** you have a rendering problem. Fixes range from enabling server-side rendering to pre-rendering key pages — a developer conversation, but a well-understood one.
- **Step 2 showed wrong information about you:** AI systems have learned something inaccurate, usually from inconsistent or outdated descriptions of your business across the web. That's an off-site cleanup job: making sure directories, review sites, and your own About page all tell the same story.

## Should I just fix the robots.txt and call it done?

Unblocking gets you *eligible*. It doesn't get you *chosen*. AI engines cite content that answers questions directly, includes concrete facts and figures, and comes from brands they see mentioned consistently across the web. Research from Princeton on generative engine optimization found that adding statistics and citing sources in your content can improve visibility in AI answers by roughly 30–40% — which is to say, the writing itself is a lever, not just the plumbing.

But plumbing first. There is no version of AI visibility that starts with a blocked crawler.

## Check all of this automatically

Everything above, you can do by hand — today, free, no signup. The honest limitation is that it's a snapshot: robots.txt files change, security defaults change, AI models change, and the check you ran in August tells you nothing about November.

That's the gap [SeeGeo's free audit](https://see-geo.com/) covers: it runs this entire check (plus about thirty others across your content structure, schema markup, and off-site mentions) in around 30 seconds, and can keep re-running it so you find out when something breaks — instead of finding out when the leads stop.

---

## Frequently asked questions

**How do I know if my website blocks GPTBot?**
Visit `yourdomain.com/robots.txt` and search the page for "GPTBot." If you find `User-agent: GPTBot` followed by `Disallow: /`, your site tells OpenAI's crawler to read nothing. Also check your CDN or firewall (like Cloudflare) settings, which can block bots regardless of robots.txt.

**Does blocking AI crawlers protect my content?**
It prevents those specific crawlers from reading your site, which some owners prefer for content-protection reasons. The trade-off in 2026 is real, though: blocked crawlers can't include you in AI answers, and AI answers are where a growing share of customer research happens. It's a legitimate business decision — it should just be a *decision*, not a default you didn't know about.

**Why does ChatGPT say wrong things about my business?**
Usually because the information about your business across the web is inconsistent, outdated, or thin. AI systems learn from many sources — your site, directories, review platforms, articles. Cleaning up and aligning those descriptions, and publishing a clear plain-language "what we do" statement on your own site, is the standard fix.

**Is being visible to ChatGPT the same as ranking on Google?**
No. They overlap (both reward crawlable sites, clear content, and authority) but diverge in important ways: AI engines lift direct answers rather than listing links, they weigh brand *mentions* even without hyperlinks, and they can be blocked independently of Google. A site can rank #1 on Google and be entirely absent from AI answers — that combination is exactly what this check detects.

**How often should I re-check my AI visibility?**
Monthly at minimum, and after any site migration, CDN change, or redesign. AI models and crawler policies change frequently enough that quarterly-or-older information in this space is considered stale even by industry reviewers.

---

# SEO vs GEO: what actually changes when your customers ask AI instead of Google
Source: https://see-geo.com/blog/seo-vs-geo-difference · Updated 2026-08-11

> SEO gets you ranked. GEO gets you cited by ChatGPT, Perplexity, and AI Overviews. What changes, what stays the same, and what's overhyped.

**Quick answer: SEO (search engine optimization) is the practice of getting your website ranked in a list of search results; GEO (generative engine optimization) is the practice of getting your business named inside the answer itself when someone asks an AI like ChatGPT, Perplexity, Gemini, or Google's AI Overviews.** They share a foundation — a crawlable site, clear content, real authority — but they reward different things at the top: SEO rewards pages that earn clicks, while GEO rewards content that can be lifted, quoted, and attributed as an answer. Most businesses in 2026 need both, and the good news is that maybe 70% of the work overlaps.

This guide explains the difference in plain English, walks through what genuinely changes, what stays the same, what's overhyped — and is honest about how young this discipline still is.

## What is GEO (generative engine optimization)?

GEO is the set of practices that make an AI system more likely to mention, cite, or recommend your business when it generates an answer. The term comes from a 2023 Princeton research paper that coined "generative engine optimization" and tested which content changes actually moved visibility in AI answers. You'll also see AEO (answer engine optimization) and "AI visibility" — functionally, all three describe the same practice, and "AI visibility" has emerged as the cleanest umbrella term.

The reason a new term exists at all: the interface changed. For twenty-five years, searching meant scanning a ranked list and choosing where to click. Increasingly, it means reading (or hearing) a single synthesized answer. A 2026 study cited by Search Engine Land found that 37% of consumers now start their research in AI tools rather than search engines, Gartner has projected a 25% decline in traditional search volume, and click-through rates on queries that trigger AI summaries have dropped by as much as 61% since mid-2024. When the list disappears, "ranking on the list" stops being the whole game. Being *inside the answer* becomes the game.

## What stays the same between SEO and GEO?

More than the hype suggests. If you've invested in real SEO, you have not wasted your money. Both disciplines reward:

**A crawlable, healthy site.** Both search engines and AI systems read your site with automated crawlers. Broken pages, blocked bots, and content that only appears after JavaScript runs hurt you in both worlds. (One divergence hides here, though — see the next section.)

**Content that actually answers what people ask.** Google has spent a decade pushing toward rewarding genuinely helpful content. AI engines take that to its logical extreme: they *are* the helpful answer, assembled from the most usable sources. Thin, keyword-stuffed pages lose in both systems.

**Authority and trust signals.** Named authors, credentials, an About page that says who you are, accurate business information, positive third-party coverage — Google calls this cluster E-E-A-T, and AI systems lean on the same signals when deciding which sources are safe to repeat.

**Structured data.** Schema markup (machine-readable labels describing what a page contains — a product, a FAQ, a local business) helps search engines build rich results and helps AI systems understand your content with confidence. It's the closest thing to a shared language across both systems.

## What actually changes with GEO?

Five things, and they're worth internalizing because they change how you write and where you spend effort.

### 1. You're competing to be the answer, not a result

A Google result wins by earning a click among ten options. An AI answer typically names one to three sources — or none. The competition is more brutal, but the reward is bigger: being *the* recommendation, delivered in a conversational sentence, often without the user comparing alternatives at all. Practically, this means each important page should open with a direct, self-contained answer to the question it targets — not three paragraphs of wind-up. AI systems lift passages; give them a passage worth lifting.

### 2. Mentions matter even without links

Classic SEO obsesses over backlinks — other sites linking to yours. AI systems learn about your brand from the entire text of the web, links or no links. A Reddit thread recommending you, a directory listing, a "best of" article that names you without linking — all of it shapes whether an AI *knows* you and what it believes about you. This is genuinely new work: auditing where your brand is mentioned, whether those descriptions are consistent, and — most actionable of all — which specific pages AI engines already cite for your target questions, so you can work on being present in those exact pages.

### 3. Formatting for extraction beats formatting for browsing

Question-shaped headings ("How much does X cost?"), short standalone passages of two to three sentences, comparison tables, bulleted specifics, visible dates — these help an AI cleanly extract and attribute your content. The Princeton GEO research quantified some of this: adding statistics and citing sources in your content improved AI visibility by roughly 30–40% in their tests, making "add real numbers and name your sources" arguably the highest-leverage writing habit in GEO.

### 4. The crawler list got longer — and some of it is blocked by default

Beyond Googlebot, you now care about [GPTBot](https://see-geo.com/bots/gptbot) (OpenAI), [ClaudeBot](https://see-geo.com/bots/claudebot) (Anthropic), [PerplexityBot](https://see-geo.com/bots/perplexitybot), and [Google-Extended](https://see-geo.com/bots/google-extended) (Google's AI crawler, separate from its search one). Many sites block these — sometimes deliberately during the 2023–24 content-scraping debates, sometimes silently via security services like Cloudflare, which turned to blocking AI bots by default. A site can be perfectly visible to Google and completely dark to every AI engine because of one setting nobody remembers choosing. [Checking this takes five minutes](https://see-geo.com/blog/is-your-website-visible-to-chatgpt) and is the single most common critical finding we see.

### 5. Measurement is a different activity entirely

SEO measurement is mature: rankings, impressions, clicks, Search Console. GEO measurement means asking the engines your customers' actual questions — "best [category] for [need]," "alternatives to [competitor]," "[your brand] reviews" — across ChatGPT, Perplexity, Gemini, and AI Overviews, and scoring whether you're mentioned, cited, or recommended, and *how you're described*. And because AI answers are non-deterministic (the same question can produce different answers on different runs), a single check is a snapshot, not a measurement. Real GEO tracking means running a panel of questions repeatedly over time and watching the trend — which is tedious by hand and is, candidly, the reason tools like ours exist.

## SEO vs GEO at a glance

| | SEO | GEO |
|---|---|---|
| **Goal** | Rank in a list of results | Be named inside the AI's answer |
| **Winning unit** | A page that earns clicks | A passage worth quoting + a brand worth naming |
| **Authority currency** | Backlinks | Mentions everywhere, linked or not |
| **Key crawlers** | Googlebot, Bingbot | GPTBot, ClaudeBot, PerplexityBot, Google-Extended |
| **Content style** | Comprehensive, browsable | Answer-first, extractable, statistic-rich |
| **Measurement** | Rankings, clicks (mature tools) | Mention/citation tracking across engines (young, repeated sampling required) |
| **Maturity** | ~25 years of evidence | ~2 years; one landmark study, evolving consensus |

## What's overhyped about GEO?

Being honest here matters more to us than sounding impressive, so:

**"SEO is dead."** No. Search volume is shifting, not vanishing, and AI engines themselves retrieve heavily from content that ranks well in traditional search. Good SEO feeds GEO. The businesses winning AI visibility right now are overwhelmingly ones with strong classic-search footprints.

**"There's a secret trick to getting cited."** Anyone selling a guaranteed AI-ranking trick is selling weather control. AI systems are non-deterministic, they update constantly, and the discipline is roughly two years old. There is one peer-reviewed landmark study (the Princeton work), a growing body of practitioner consensus, and a lot of confident guessing. The honest pitch is: there are clearly better and worse practices, the better ones measurably help, and everything should be re-tested as models change — the industry's own reviewers warn that comparison data in this space goes stale within a quarter.

**"You need to choose between optimizing for Google and optimizing for AI."** For roughly 70% of the work — site health, clear answers, real authority, structured data — they're the same work. The genuinely GEO-specific slice (AI-crawler access, extraction-friendly formatting, mention-building, answer tracking) is additive, not a fork in the road.

## So what should a small business actually do?

In order of leverage:

1. **Confirm access.** Check that AI crawlers aren't blocked in your robots.txt or by your CDN, and that your content is readable without JavaScript. Five minutes; everything else depends on it.
2. **Make your best pages answer-first.** Open with the direct answer, use question-shaped headings, add a real statistic with a named source, put comparisons in tables, show dates and authors.
3. **Add or fix structured data** — especially FAQ and HowTo markup on pages shaped like questions and instructions.
4. **Audit your mentions.** Ask the AI engines your customers' questions, note which sources get cited, and work on appearing in those specific places. Fix inconsistent descriptions of your business across the web.
5. **Measure on a schedule, not once.** Re-run your question panel monthly and watch the trend.

Or let software do steps 1–5 continuously: [SeeGeo's free audit](https://see-geo.com/) runs the access, formatting, and structured-data checks in about 30 seconds, and the paid plan keeps a live question panel running across the major AI engines so you see your visibility as a trend line instead of a guess.

---

## Frequently asked questions

**What does GEO stand for in marketing?**
Generative engine optimization — the practice of improving how often and how favorably AI systems like ChatGPT, Perplexity, Gemini, and Google's AI Overviews mention or cite your business in their answers. It is unrelated to "geo" as in geographic targeting, which is a different, older marketing term.

**Is GEO the same as AEO?**
Functionally yes. GEO (generative engine optimization) and AEO (answer engine optimization) describe the same practice; "AI visibility" is the cleanest umbrella term and the one gaining consensus.

**Do I still need SEO if I do GEO?**
Yes. AI engines draw heavily on content that performs well in traditional search, and search itself still carries enormous volume. The two overlap on most fundamentals; GEO adds a layer rather than replacing the stack.

**How do I show up in ChatGPT recommendations?**
Ensure AI crawlers can access your site, publish answer-first content with concrete facts and cited sources, maintain consistent business information across the web, and build presence on the third-party pages AI engines already cite in your category. Then measure repeatedly — single checks are unreliable because AI answers vary between runs.

**How long does GEO take to work?**
Access fixes (unblocking crawlers) can reflect in retrieval-based engines like Perplexity within days to weeks. Changes that depend on model training data or broad web presence move slower — think months. Anyone promising precise timelines is overpromising; the honest answer is that fast wins exist (access, formatting, being added to already-cited pages) and slow wins compound (brand mentions, authority).

**Is GEO worth it for a small local business?**
Increasingly yes, because "best [category] near me"-style questions are exactly what people ask AI assistants. For local businesses the fundamentals are consistent name/address/phone information everywhere, a clear plain-language description of what you do, review presence, and an unblocked site. The overlap with good local SEO is nearly total — which means most of the work pays off twice.

---

# Should you block AI crawlers? What GPTBot, ClaudeBot & friends actually do
Source: https://see-geo.com/blog/should-you-block-ai-crawlers · Updated 2026-08-10

> Blocking AI bots feels protective, but it removes you from the answers your customers read. What each crawler does, and the two common mistakes.

Short answer: **for most small businesses, no** — blocking AI crawlers removes you from the AI answers your customers increasingly rely on, and the two most common blocks we see in audits are misunderstandings, not decisions.

The longer answer depends on which bot, because "AI crawler" covers three very different jobs.

## What are the three kinds of AI crawler?

**Training crawlers** (GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot) collect pages so future models know your business exists. Blocking them is a real trade-off: your content stays out of training data, and future models learn about your competitors instead of you.

**Search-index crawlers** (OAI-SearchBot, PerplexityBot) build the live indexes behind AI search. Block these and the AI *cannot cite you* even when you're the best answer.

**On-demand fetchers** (ChatGPT-User, Claude-User, Perplexity-User) fetch your page live when a user asks about you mid-conversation. Blocking them means the AI answers questions about your business without being able to check your site.

## Which blocks are usually mistakes?

Two patterns come up constantly in SeeGeo audits ([robots.txt](https://www.rfc-editor.org/rfc/rfc9309) rules per RFC 9309):

**The blanket block.** A `User-agent: * / Disallow: /` left over from a staging site, or added "for security." It doesn't just block AI — it blocks Googlebot. Everything else on the site stops mattering until it's removed.

**The Google-Extended confusion.** Blocking Google-Extended does **not** remove you from Google Search or AI Overviews — those run on Googlebot. It only opts you out of Gemini model training. The reverse mistake is worse: some owners block Bingbot "because we don't care about Bing," not realizing AI assistants that draw on Bing's index — including parts of ChatGPT's search — lose access too.

## When is blocking the right call?

Legitimate reasons exist. Publishers whose content *is* the product may rationally block training crawlers. Sites with user data behind logins should block everything from private paths. And any business can decide the principle matters more than the visibility. The key is making it a **decision, not an accident**: know which bots you're blocking, what each one feeds, and what you're giving up.

## How do you check what you're blocking right now?

Three places to look, because robots.txt alone isn't the whole story:

1. **robots.txt** — read the actual rules per bot (or run a [SeeGeo audit](https://see-geo.com/), which evaluates all 16 major crawlers individually).
2. **Your CDN** — Cloudflare blocks AI crawlers **by default** for many accounts now. Your robots.txt can say "welcome" while the firewall says no.
3. **Your rendering** — most AI crawlers don't run JavaScript. If your content only appears client-side, you're "blocking" every AI bot without a single rule.

The uncomfortable truth: plenty of businesses that never chose to block anything are invisible to AI anyway. That's why we built the audit to check all three layers.

---

# What is GEO? Generative engine optimization, in plain English
Source: https://see-geo.com/blog/what-is-geo · Updated 2026-08-10

> GEO is making your business visible inside AI-written answers. What it involves, how it differs from SEO, and what the research actually supports.

GEO — generative engine optimization — is the practice of making your business show up **inside AI-written answers**: when someone asks ChatGPT, Claude, Gemini, or Google's AI Overviews for a recommendation, GEO is what determines whether you're named, cited, or skipped.

That's the whole definition. The rest of this post is what it means in practice.

## How is GEO different from SEO?

SEO optimizes for a **list**: ten blue links, ranked, and your job is to be near the top. GEO optimizes for an **answer**: the AI writes two or three sentences, names two or three businesses, and everyone else is invisible.

The overlap between the two is large — and that's good news. Google's AI Overviews are built on the same Googlebot crawl as regular search. Sites that are fast, structured, and clearly written do better at both. Where they diverge:

- **Citations replace rankings.** In an AI answer there is no position #7. You're either one of the few sources the answer is built from, or you're absent.
- **Mentions matter without links.** AI models learn about businesses from the web at large — a Reddit thread or a "best bakeries" listicle that names you (without linking) still teaches the model you exist.
- **Machines are the reader.** Most AI crawlers don't run JavaScript and don't infer. If your site never plainly says what you do and where, the AI can't confidently recommend you — so it doesn't.

## What actually works?

The best evidence available is the Princeton GEO study ([arXiv:2311.09735](https://arxiv.org/abs/2311.09735), presented at KDD 2024), which tested nine content strategies across thousands of queries. The two that consistently won: **adding concrete statistics** and **citing sources** — improvements on the order of 30–40% in how often content appeared in generative answers. Quotation-adding and fluency edits also helped; keyword stuffing did not.

Treat those numbers as research findings, not guarantees — the study measured specific engines at a specific time, and AI systems change fast. But as a priority order it's the best-evidenced starting point there is:

1. **Be readable**: no JavaScript-only content, no accidental crawler blocks, no CDN silently turning AI bots away.
2. **Be clear**: say what you are, where you operate, and what you cost, in plain sentences a machine can lift.
3. **Be referenced**: get into the directories, review platforms, and listicles AI already cites for your category.
4. **Be measurable**: ask the engines your customers' questions on a schedule and watch the trend — AI answers don't show up in your analytics.

## Should you hire someone for GEO?

Honest answer: the discipline is young. There is no GEO equivalent of twenty years of settled SEO practice, and anyone promising guaranteed AI rankings is overselling. Most of the work — readable site, clear entity, real off-site presence — is well within reach of a small team with the right checklist. That's the gap SeeGeo exists to close: audit what's broken, fix what matters, measure whether it worked.

---

# AI crawlers explained: every bot that matters
Source: https://see-geo.com/bots · Updated 2026-09-08

Every major AI and search crawler: what it does with your content, whether to block it, and the exact robots.txt lines — in plain language.

## Pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot) — What GPTBot (OpenAI) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot) — What OAI-SearchBot (OpenAI) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is ChatGPT-User?](https://see-geo.com/bots/chatgpt-user) — What ChatGPT-User (OpenAI) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot) — What ClaudeBot (Anthropic) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Claude-User?](https://see-geo.com/bots/claude-user) — What Claude-User (Anthropic) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is PerplexityBot?](https://see-geo.com/bots/perplexitybot) — What PerplexityBot (Perplexity) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Perplexity-User?](https://see-geo.com/bots/perplexity-user) — What Perplexity-User (Perplexity) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Google-Extended?](https://see-geo.com/bots/google-extended) — What Google-Extended (Google) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is CCBot?](https://see-geo.com/bots/ccbot) — What CCBot (Common Crawl) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Bytespider?](https://see-geo.com/bots/bytespider) — What Bytespider (ByteDance) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Amazonbot?](https://see-geo.com/bots/amazonbot) — What Amazonbot (Amazon) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Applebot-Extended?](https://see-geo.com/bots/applebot-extended) — What Applebot-Extended (Apple) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Meta-ExternalAgent?](https://see-geo.com/bots/meta-externalagent) — What Meta-ExternalAgent (Meta) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is DuckAssistBot?](https://see-geo.com/bots/duckassistbot) — What DuckAssistBot (DuckDuckGo) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Googlebot?](https://see-geo.com/bots/googlebot) — What Googlebot (Google) does with your content, whether to block it, and the exact robots.txt lines — in plain language.
- [What is Bingbot?](https://see-geo.com/bots/bingbot) — What Bingbot (Microsoft) does with your content, whether to block it, and the exact robots.txt lines — in plain language.

---

# What is GPTBot?
Source: https://see-geo.com/bots/gptbot · Updated 2026-09-08

GPTBot is OpenAI's AI training crawler. Collects pages to train ChatGPT's underlying models.

## What does GPTBot do with your content?

Collects pages to train ChatGPT's underlying models. Like most AI crawlers, GPTBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block GPTBot?

This is a genuine trade-off. Allowing GPTBot lets future OpenAI models learn your business exists — useful when customers ask those models for recommendations. Blocking it keeps your content out of training data at the cost of that future visibility. Neither answer is wrong; it depends on what you value more.

## How do you allow or block GPTBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows GPTBot everywhere; the second blocks it completely.

```
# Allow GPTBot
User-agent: GPTBot
Allow: /

# Block GPTBot
User-agent: GPTBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want GPTBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether GPTBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming GPTBot (or the * group it falls back to), then request your homepage with GPTBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does GPTBot compare to similar crawlers?

GPTBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| GPTBot | OpenAI | AI training crawler | No |
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| ChatGPT-User | OpenAI | on-demand AI fetcher | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is GPTBot?

GPTBot is OpenAI's AI training crawler. Collects pages to train ChatGPT's underlying models.

### Should I block GPTBot?

This is a genuine trade-off. Allowing GPTBot lets future OpenAI models learn your business exists — useful when customers ask those models for recommendations.

### Does GPTBot run JavaScript?

No. GPTBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can GPTBot read my website?

Only if two things are true: your robots.txt does not disallow GPTBot, and your server or CDN actually serves the page when GPTBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as GPTBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot)
- [What is ChatGPT-User?](https://see-geo.com/bots/chatgpt-user)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)

---

# What is OAI-SearchBot?
Source: https://see-geo.com/bots/oai-searchbot · Updated 2026-09-08

OAI-SearchBot is OpenAI's AI search index crawler. Builds the index behind ChatGPT search — this is the bot that decides if ChatGPT can cite you.

## What does OAI-SearchBot do with your content?

Builds the index behind ChatGPT search — this is the bot that decides if ChatGPT can cite you. Like most AI crawlers, OAI-SearchBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block OAI-SearchBot?

Blocking OAI-SearchBot makes your business invisible to OpenAI's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep OpenAI's AI away from your content.

## How do you allow or block OAI-SearchBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows OAI-SearchBot everywhere; the second blocks it completely.

```
# Allow OAI-SearchBot
User-agent: OAI-SearchBot
Allow: /

# Block OAI-SearchBot
User-agent: OAI-SearchBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want OAI-SearchBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether OAI-SearchBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming OAI-SearchBot (or the * group it falls back to), then request your homepage with OAI-SearchBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does OAI-SearchBot compare to similar crawlers?

OAI-SearchBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| GPTBot | OpenAI | AI training crawler | No |
| ChatGPT-User | OpenAI | on-demand AI fetcher | No |
| PerplexityBot | Perplexity | AI search index crawler | No |
| Amazonbot | Amazon | AI search index crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is OAI-SearchBot?

OAI-SearchBot is OpenAI's AI search index crawler. Builds the index behind ChatGPT search — this is the bot that decides if ChatGPT can cite you.

### Should I block OAI-SearchBot?

Blocking OAI-SearchBot makes your business invisible to OpenAI's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep OpenAI's AI away from your content..

### Does OAI-SearchBot run JavaScript?

No. OAI-SearchBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can OAI-SearchBot read my website?

Only if two things are true: your robots.txt does not disallow OAI-SearchBot, and your server or CDN actually serves the page when OAI-SearchBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as OAI-SearchBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ChatGPT-User?](https://see-geo.com/bots/chatgpt-user)
- [What is PerplexityBot?](https://see-geo.com/bots/perplexitybot)

---

# What is ChatGPT-User?
Source: https://see-geo.com/bots/chatgpt-user · Updated 2026-09-08

ChatGPT-User is OpenAI's on-demand AI fetcher. Fetches your page live when a ChatGPT user asks about you in a conversation.

## What does ChatGPT-User do with your content?

Fetches your page live when a ChatGPT user asks about you in a conversation. Like most AI crawlers, ChatGPT-User reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block ChatGPT-User?

Blocking ChatGPT-User makes your business invisible to OpenAI's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep OpenAI's AI away from your content.

## How do you allow or block ChatGPT-User in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows ChatGPT-User everywhere; the second blocks it completely.

```
# Allow ChatGPT-User
User-agent: ChatGPT-User
Allow: /

# Block ChatGPT-User
User-agent: ChatGPT-User
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want ChatGPT-User to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether ChatGPT-User can read my site right now?

Two checks, both needed: read your robots.txt for a group naming ChatGPT-User (or the * group it falls back to), then request your homepage with ChatGPT-User's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does ChatGPT-User compare to similar crawlers?

ChatGPT-User is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| ChatGPT-User | OpenAI | on-demand AI fetcher | No |
| GPTBot | OpenAI | AI training crawler | No |
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| Claude-User | Anthropic | on-demand AI fetcher | No |
| Perplexity-User | Perplexity | on-demand AI fetcher | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is ChatGPT-User?

ChatGPT-User is OpenAI's on-demand AI fetcher. Fetches your page live when a ChatGPT user asks about you in a conversation.

### Should I block ChatGPT-User?

Blocking ChatGPT-User makes your business invisible to OpenAI's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep OpenAI's AI away from your content..

### Does ChatGPT-User run JavaScript?

No. ChatGPT-User, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can ChatGPT-User read my website?

Only if two things are true: your robots.txt does not disallow ChatGPT-User, and your server or CDN actually serves the page when ChatGPT-User asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as ChatGPT-User — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot)
- [What is Claude-User?](https://see-geo.com/bots/claude-user)

---

# What is ClaudeBot?
Source: https://see-geo.com/bots/claudebot · Updated 2026-09-08

ClaudeBot is Anthropic's AI training crawler. Crawls pages for Claude's models and search index.

## What does ClaudeBot do with your content?

Crawls pages for Claude's models and search index. Like most AI crawlers, ClaudeBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block ClaudeBot?

This is a genuine trade-off. Allowing ClaudeBot lets future Anthropic models learn your business exists — useful when customers ask those models for recommendations. Blocking it keeps your content out of training data at the cost of that future visibility. Neither answer is wrong; it depends on what you value more.

## How do you allow or block ClaudeBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows ClaudeBot everywhere; the second blocks it completely.

```
# Allow ClaudeBot
User-agent: ClaudeBot
Allow: /

# Block ClaudeBot
User-agent: ClaudeBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want ClaudeBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether ClaudeBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming ClaudeBot (or the * group it falls back to), then request your homepage with ClaudeBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does ClaudeBot compare to similar crawlers?

ClaudeBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| ClaudeBot | Anthropic | AI training crawler | No |
| Claude-User | Anthropic | on-demand AI fetcher | No |
| GPTBot | OpenAI | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |
| CCBot | Common Crawl | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is ClaudeBot?

ClaudeBot is Anthropic's AI training crawler. Crawls pages for Claude's models and search index.

### Should I block ClaudeBot?

This is a genuine trade-off. Allowing ClaudeBot lets future Anthropic models learn your business exists — useful when customers ask those models for recommendations.

### Does ClaudeBot run JavaScript?

No. ClaudeBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can ClaudeBot read my website?

Only if two things are true: your robots.txt does not disallow ClaudeBot, and your server or CDN actually serves the page when ClaudeBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as ClaudeBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is Claude-User?](https://see-geo.com/bots/claude-user)
- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is Google-Extended?](https://see-geo.com/bots/google-extended)

---

# What is Claude-User?
Source: https://see-geo.com/bots/claude-user · Updated 2026-09-08

Claude-User is Anthropic's on-demand AI fetcher. Fetches your page live during a Claude conversation.

## What does Claude-User do with your content?

Fetches your page live during a Claude conversation. Like most AI crawlers, Claude-User reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Claude-User?

Blocking Claude-User makes your business invisible to Anthropic's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Anthropic's AI away from your content.

## How do you allow or block Claude-User in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Claude-User everywhere; the second blocks it completely.

```
# Allow Claude-User
User-agent: Claude-User
Allow: /

# Block Claude-User
User-agent: Claude-User
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Claude-User to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Claude-User can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Claude-User (or the * group it falls back to), then request your homepage with Claude-User's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Claude-User compare to similar crawlers?

Claude-User is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Claude-User | Anthropic | on-demand AI fetcher | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| ChatGPT-User | OpenAI | on-demand AI fetcher | No |
| Perplexity-User | Perplexity | on-demand AI fetcher | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Claude-User?

Claude-User is Anthropic's on-demand AI fetcher. Fetches your page live during a Claude conversation.

### Should I block Claude-User?

Blocking Claude-User makes your business invisible to Anthropic's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Anthropic's AI away from your content..

### Does Claude-User run JavaScript?

No. Claude-User, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Claude-User read my website?

Only if two things are true: your robots.txt does not disallow Claude-User, and your server or CDN actually serves the page when Claude-User asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Claude-User — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)
- [What is ChatGPT-User?](https://see-geo.com/bots/chatgpt-user)
- [What is Perplexity-User?](https://see-geo.com/bots/perplexity-user)

---

# What is PerplexityBot?
Source: https://see-geo.com/bots/perplexitybot · Updated 2026-09-08

PerplexityBot is Perplexity's AI search index crawler. Indexes pages for Perplexity's answer engine — Perplexity cites sources on almost every answer.

## What does PerplexityBot do with your content?

Indexes pages for Perplexity's answer engine — Perplexity cites sources on almost every answer. Like most AI crawlers, PerplexityBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block PerplexityBot?

Blocking PerplexityBot makes your business invisible to Perplexity's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Perplexity's AI away from your content.

## How do you allow or block PerplexityBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows PerplexityBot everywhere; the second blocks it completely.

```
# Allow PerplexityBot
User-agent: PerplexityBot
Allow: /

# Block PerplexityBot
User-agent: PerplexityBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want PerplexityBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether PerplexityBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming PerplexityBot (or the * group it falls back to), then request your homepage with PerplexityBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does PerplexityBot compare to similar crawlers?

PerplexityBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| PerplexityBot | Perplexity | AI search index crawler | No |
| Perplexity-User | Perplexity | on-demand AI fetcher | No |
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| Amazonbot | Amazon | AI search index crawler | No |
| DuckAssistBot | DuckDuckGo | AI search index crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is PerplexityBot?

PerplexityBot is Perplexity's AI search index crawler. Indexes pages for Perplexity's answer engine — Perplexity cites sources on almost every answer.

### Should I block PerplexityBot?

Blocking PerplexityBot makes your business invisible to Perplexity's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Perplexity's AI away from your content..

### Does PerplexityBot run JavaScript?

No. PerplexityBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can PerplexityBot read my website?

Only if two things are true: your robots.txt does not disallow PerplexityBot, and your server or CDN actually serves the page when PerplexityBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as PerplexityBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is Perplexity-User?](https://see-geo.com/bots/perplexity-user)
- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot)
- [What is Amazonbot?](https://see-geo.com/bots/amazonbot)

---

# What is Perplexity-User?
Source: https://see-geo.com/bots/perplexity-user · Updated 2026-09-08

Perplexity-User is Perplexity's on-demand AI fetcher. Fetches your page live when a Perplexity user asks.

## What does Perplexity-User do with your content?

Fetches your page live when a Perplexity user asks. Like most AI crawlers, Perplexity-User reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Perplexity-User?

Blocking Perplexity-User makes your business invisible to Perplexity's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Perplexity's AI away from your content.

## How do you allow or block Perplexity-User in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Perplexity-User everywhere; the second blocks it completely.

```
# Allow Perplexity-User
User-agent: Perplexity-User
Allow: /

# Block Perplexity-User
User-agent: Perplexity-User
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Perplexity-User to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Perplexity-User can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Perplexity-User (or the * group it falls back to), then request your homepage with Perplexity-User's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Perplexity-User compare to similar crawlers?

Perplexity-User is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Perplexity-User | Perplexity | on-demand AI fetcher | No |
| PerplexityBot | Perplexity | AI search index crawler | No |
| ChatGPT-User | OpenAI | on-demand AI fetcher | No |
| Claude-User | Anthropic | on-demand AI fetcher | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Perplexity-User?

Perplexity-User is Perplexity's on-demand AI fetcher. Fetches your page live when a Perplexity user asks.

### Should I block Perplexity-User?

Blocking Perplexity-User makes your business invisible to Perplexity's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Perplexity's AI away from your content..

### Does Perplexity-User run JavaScript?

No. Perplexity-User, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Perplexity-User read my website?

Only if two things are true: your robots.txt does not disallow Perplexity-User, and your server or CDN actually serves the page when Perplexity-User asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Perplexity-User — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is PerplexityBot?](https://see-geo.com/bots/perplexitybot)
- [What is ChatGPT-User?](https://see-geo.com/bots/chatgpt-user)
- [What is Claude-User?](https://see-geo.com/bots/claude-user)

---

# What is Google-Extended?
Source: https://see-geo.com/bots/google-extended · Updated 2026-09-08

Google-Extended is Google's AI training crawler. Controls whether Google may train Gemini on your content. Blocking it does NOT affect Google Search or AI Overviews.

## What does Google-Extended do with your content?

Controls whether Google may train Gemini on your content. Blocking it does NOT affect Google Search or AI Overviews. Like most AI crawlers, Google-Extended reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Google-Extended?

Blocking Google-Extended only opts your content out of Gemini model training. It does NOT remove you from Google Search or from AI Overviews — both of those run on Googlebot. If your goal is staying visible in Google while opting out of model training, blocking Google-Extended is a reasonable, low-risk choice.

## How do you allow or block Google-Extended in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Google-Extended everywhere; the second blocks it completely.

```
# Allow Google-Extended
User-agent: Google-Extended
Allow: /

# Block Google-Extended
User-agent: Google-Extended
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Google-Extended to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Google-Extended can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Google-Extended (or the * group it falls back to), then request your homepage with Google-Extended's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Google-Extended compare to similar crawlers?

Google-Extended is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Google-Extended | Google | AI training crawler | No |
| Googlebot | Google | search engine crawler | Delayed rendering |
| GPTBot | OpenAI | AI training crawler | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| CCBot | Common Crawl | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Google-Extended?

Google-Extended is Google's AI training crawler. Controls whether Google may train Gemini on your content. Blocking it does NOT affect Google Search or AI Overviews.

### Should I block Google-Extended?

Blocking Google-Extended only opts your content out of Gemini model training. It does NOT remove you from Google Search or from AI Overviews — both of those run on Googlebot.

### Does Google-Extended run JavaScript?

No. Google-Extended, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Google-Extended read my website?

Only if two things are true: your robots.txt does not disallow Google-Extended, and your server or CDN actually serves the page when Google-Extended asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Google-Extended — the free check on see-geo.com/ai-crawler-check does exactly that.

### Does blocking Google-Extended affect Google Search rankings?

No. Google-Extended only controls Gemini model training. Google Search and AI Overviews use Googlebot, so blocking Google-Extended has no effect on either.

## Related pages

- [What is Googlebot?](https://see-geo.com/bots/googlebot)
- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)

---

# What is CCBot?
Source: https://see-geo.com/bots/ccbot · Updated 2026-09-08

CCBot is Common Crawl's AI training crawler. Builds a public web archive that many AI companies train models on. Blocking it quietly removes you from future models' knowledge.

## What does CCBot do with your content?

Builds a public web archive that many AI companies train models on. Blocking it quietly removes you from future models' knowledge. Like most AI crawlers, CCBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block CCBot?

CCBot feeds Common Crawl, a public archive many AI labs train on. Blocking it quietly removes you from the training data of future models — models that customers may ask for recommendations years from now. The trade-off is real: allow it for AI visibility, block it if keeping your content out of training data matters more to you.

## How do you allow or block CCBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows CCBot everywhere; the second blocks it completely.

```
# Allow CCBot
User-agent: CCBot
Allow: /

# Block CCBot
User-agent: CCBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want CCBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether CCBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming CCBot (or the * group it falls back to), then request your homepage with CCBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does CCBot compare to similar crawlers?

CCBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| CCBot | Common Crawl | AI training crawler | No |
| GPTBot | OpenAI | AI training crawler | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |
| Bytespider | ByteDance | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is CCBot?

CCBot is Common Crawl's AI training crawler. Builds a public web archive that many AI companies train models on. Blocking it quietly removes you from future models' knowledge.

### Should I block CCBot?

CCBot feeds Common Crawl, a public archive many AI labs train on. Blocking it quietly removes you from the training data of future models — models that customers may ask for recommendations years from now.

### Does CCBot run JavaScript?

No. CCBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can CCBot read my website?

Only if two things are true: your robots.txt does not disallow CCBot, and your server or CDN actually serves the page when CCBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as CCBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)
- [What is Google-Extended?](https://see-geo.com/bots/google-extended)

---

# What is Bytespider?
Source: https://see-geo.com/bots/bytespider · Updated 2026-09-08

Bytespider is ByteDance's AI training crawler. Trains ByteDance's models (Doubao).

## What does Bytespider do with your content?

Trains ByteDance's models (Doubao). Like most AI crawlers, Bytespider reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Bytespider?

This is a genuine trade-off. Allowing Bytespider lets future ByteDance models learn your business exists — useful when customers ask those models for recommendations. Blocking it keeps your content out of training data at the cost of that future visibility. Neither answer is wrong; it depends on what you value more.

## How do you allow or block Bytespider in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Bytespider everywhere; the second blocks it completely.

```
# Allow Bytespider
User-agent: Bytespider
Allow: /

# Block Bytespider
User-agent: Bytespider
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Bytespider to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Bytespider can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Bytespider (or the * group it falls back to), then request your homepage with Bytespider's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Bytespider compare to similar crawlers?

Bytespider is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Bytespider | ByteDance | AI training crawler | No |
| GPTBot | OpenAI | AI training crawler | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |
| CCBot | Common Crawl | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Bytespider?

Bytespider is ByteDance's AI training crawler. Trains ByteDance's models (Doubao).

### Should I block Bytespider?

This is a genuine trade-off. Allowing Bytespider lets future ByteDance models learn your business exists — useful when customers ask those models for recommendations.

### Does Bytespider run JavaScript?

No. Bytespider, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Bytespider read my website?

Only if two things are true: your robots.txt does not disallow Bytespider, and your server or CDN actually serves the page when Bytespider asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Bytespider — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)
- [What is Google-Extended?](https://see-geo.com/bots/google-extended)

---

# What is Amazonbot?
Source: https://see-geo.com/bots/amazonbot · Updated 2026-09-08

Amazonbot is Amazon's AI search index crawler. Feeds Alexa and Amazon's AI answers.

## What does Amazonbot do with your content?

Feeds Alexa and Amazon's AI answers. Like most AI crawlers, Amazonbot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Amazonbot?

Blocking Amazonbot makes your business invisible to Amazon's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Amazon's AI away from your content.

## How do you allow or block Amazonbot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Amazonbot everywhere; the second blocks it completely.

```
# Allow Amazonbot
User-agent: Amazonbot
Allow: /

# Block Amazonbot
User-agent: Amazonbot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Amazonbot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Amazonbot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Amazonbot (or the * group it falls back to), then request your homepage with Amazonbot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Amazonbot compare to similar crawlers?

Amazonbot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Amazonbot | Amazon | AI search index crawler | No |
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| PerplexityBot | Perplexity | AI search index crawler | No |
| DuckAssistBot | DuckDuckGo | AI search index crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Amazonbot?

Amazonbot is Amazon's AI search index crawler. Feeds Alexa and Amazon's AI answers.

### Should I block Amazonbot?

Blocking Amazonbot makes your business invisible to Amazon's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep Amazon's AI away from your content..

### Does Amazonbot run JavaScript?

No. Amazonbot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Amazonbot read my website?

Only if two things are true: your robots.txt does not disallow Amazonbot, and your server or CDN actually serves the page when Amazonbot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Amazonbot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot)
- [What is PerplexityBot?](https://see-geo.com/bots/perplexitybot)
- [What is DuckAssistBot?](https://see-geo.com/bots/duckassistbot)

---

# What is Applebot-Extended?
Source: https://see-geo.com/bots/applebot-extended · Updated 2026-09-08

Applebot-Extended is Apple's AI training crawler. Controls whether Apple may use your content for Apple Intelligence.

## What does Applebot-Extended do with your content?

Controls whether Apple may use your content for Apple Intelligence. Like most AI crawlers, Applebot-Extended reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Applebot-Extended?

This is a genuine trade-off. Allowing Applebot-Extended lets future Apple models learn your business exists — useful when customers ask those models for recommendations. Blocking it keeps your content out of training data at the cost of that future visibility. Neither answer is wrong; it depends on what you value more.

## How do you allow or block Applebot-Extended in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Applebot-Extended everywhere; the second blocks it completely.

```
# Allow Applebot-Extended
User-agent: Applebot-Extended
Allow: /

# Block Applebot-Extended
User-agent: Applebot-Extended
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Applebot-Extended to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Applebot-Extended can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Applebot-Extended (or the * group it falls back to), then request your homepage with Applebot-Extended's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Applebot-Extended compare to similar crawlers?

Applebot-Extended is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Applebot-Extended | Apple | AI training crawler | No |
| GPTBot | OpenAI | AI training crawler | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |
| CCBot | Common Crawl | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Applebot-Extended?

Applebot-Extended is Apple's AI training crawler. Controls whether Apple may use your content for Apple Intelligence.

### Should I block Applebot-Extended?

This is a genuine trade-off. Allowing Applebot-Extended lets future Apple models learn your business exists — useful when customers ask those models for recommendations.

### Does Applebot-Extended run JavaScript?

No. Applebot-Extended, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Applebot-Extended read my website?

Only if two things are true: your robots.txt does not disallow Applebot-Extended, and your server or CDN actually serves the page when Applebot-Extended asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Applebot-Extended — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)
- [What is Google-Extended?](https://see-geo.com/bots/google-extended)

---

# What is Meta-ExternalAgent?
Source: https://see-geo.com/bots/meta-externalagent · Updated 2026-09-08

Meta-ExternalAgent is Meta's AI training crawler. Collects pages to train Meta's Llama models.

## What does Meta-ExternalAgent do with your content?

Collects pages to train Meta's Llama models. Like most AI crawlers, Meta-ExternalAgent reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block Meta-ExternalAgent?

This is a genuine trade-off. Allowing Meta-ExternalAgent lets future Meta models learn your business exists — useful when customers ask those models for recommendations. Blocking it keeps your content out of training data at the cost of that future visibility. Neither answer is wrong; it depends on what you value more.

## How do you allow or block Meta-ExternalAgent in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Meta-ExternalAgent everywhere; the second blocks it completely.

```
# Allow Meta-ExternalAgent
User-agent: Meta-ExternalAgent
Allow: /

# Block Meta-ExternalAgent
User-agent: Meta-ExternalAgent
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Meta-ExternalAgent to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Meta-ExternalAgent can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Meta-ExternalAgent (or the * group it falls back to), then request your homepage with Meta-ExternalAgent's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Meta-ExternalAgent compare to similar crawlers?

Meta-ExternalAgent is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Meta-ExternalAgent | Meta | AI training crawler | No |
| GPTBot | OpenAI | AI training crawler | No |
| ClaudeBot | Anthropic | AI training crawler | No |
| Google-Extended | Google | AI training crawler | No |
| CCBot | Common Crawl | AI training crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Meta-ExternalAgent?

Meta-ExternalAgent is Meta's AI training crawler. Collects pages to train Meta's Llama models.

### Should I block Meta-ExternalAgent?

This is a genuine trade-off. Allowing Meta-ExternalAgent lets future Meta models learn your business exists — useful when customers ask those models for recommendations.

### Does Meta-ExternalAgent run JavaScript?

No. Meta-ExternalAgent, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can Meta-ExternalAgent read my website?

Only if two things are true: your robots.txt does not disallow Meta-ExternalAgent, and your server or CDN actually serves the page when Meta-ExternalAgent asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Meta-ExternalAgent — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is GPTBot?](https://see-geo.com/bots/gptbot)
- [What is ClaudeBot?](https://see-geo.com/bots/claudebot)
- [What is Google-Extended?](https://see-geo.com/bots/google-extended)

---

# What is DuckAssistBot?
Source: https://see-geo.com/bots/duckassistbot · Updated 2026-09-08

DuckAssistBot is DuckDuckGo's AI search index crawler. Powers DuckDuckGo's AI-assisted answers.

## What does DuckAssistBot do with your content?

Powers DuckDuckGo's AI-assisted answers. Like most AI crawlers, DuckAssistBot reads your pages as plain HTML — it does not run JavaScript — so content that only appears after scripts execute is invisible to it.

## Should you block DuckAssistBot?

Blocking DuckAssistBot makes your business invisible to DuckDuckGo's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep DuckDuckGo's AI away from your content.

## How do you allow or block DuckAssistBot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows DuckAssistBot everywhere; the second blocks it completely.

```
# Allow DuckAssistBot
User-agent: DuckAssistBot
Allow: /

# Block DuckAssistBot
User-agent: DuckAssistBot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want DuckAssistBot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether DuckAssistBot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming DuckAssistBot (or the * group it falls back to), then request your homepage with DuckAssistBot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does DuckAssistBot compare to similar crawlers?

DuckAssistBot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| DuckAssistBot | DuckDuckGo | AI search index crawler | No |
| OAI-SearchBot | OpenAI | AI search index crawler | No |
| PerplexityBot | Perplexity | AI search index crawler | No |
| Amazonbot | Amazon | AI search index crawler | No |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is DuckAssistBot?

DuckAssistBot is DuckDuckGo's AI search index crawler. Powers DuckDuckGo's AI-assisted answers.

### Should I block DuckAssistBot?

Blocking DuckAssistBot makes your business invisible to DuckDuckGo's AI when customers ask it for recommendations. If you want to be found, allow it; block it only if you have a specific reason to keep DuckDuckGo's AI away from your content..

### Does DuckAssistBot run JavaScript?

No. DuckAssistBot, like most AI crawlers, reads raw HTML without executing JavaScript. If your content only renders client-side, it is effectively invisible to it.

### Can DuckAssistBot read my website?

Only if two things are true: your robots.txt does not disallow DuckAssistBot, and your server or CDN actually serves the page when DuckAssistBot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as DuckAssistBot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is OAI-SearchBot?](https://see-geo.com/bots/oai-searchbot)
- [What is PerplexityBot?](https://see-geo.com/bots/perplexitybot)
- [What is Amazonbot?](https://see-geo.com/bots/amazonbot)

---

# What is Googlebot?
Source: https://see-geo.com/bots/googlebot · Updated 2026-09-08

Googlebot is Google's search engine crawler. Google's main crawler. Powers Google Search AND the AI Overviews box at the top of results.

## What does Googlebot do with your content?

Google's main crawler. Powers Google Search AND the AI Overviews box at the top of results. It follows the rules you publish in robots.txt, a small text file at the root of your site that tells automated visitors what they may read (defined by RFC 9309, the Robots Exclusion Protocol standard).

## Should you block Googlebot?

Blocking Googlebot removes you from Google Search AND from AI Overviews, which are built on Googlebot's index. For any business that wants customers to find it, blocking Googlebot is almost never the right call.

## How do you allow or block Googlebot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Googlebot everywhere; the second blocks it completely.

```
# Allow Googlebot
User-agent: Googlebot
Allow: /

# Block Googlebot
User-agent: Googlebot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Googlebot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Googlebot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Googlebot (or the * group it falls back to), then request your homepage with Googlebot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Googlebot compare to similar crawlers?

Googlebot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Googlebot | Google | search engine crawler | Delayed rendering |
| Google-Extended | Google | AI training crawler | No |
| Bingbot | Microsoft | search engine crawler | Delayed rendering |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Googlebot?

Googlebot is Google's search engine crawler. Google's main crawler. Powers Google Search AND the AI Overviews box at the top of results.

### Should I block Googlebot?

Blocking Googlebot removes you from Google Search AND from AI Overviews, which are built on Googlebot's index. For any business that wants customers to find it, blocking Googlebot is almost never the right call..

### Does Googlebot run JavaScript?

Googlebot can render JavaScript for indexing, but rendering is delayed and imperfect — server-rendered content is always the safer path.

### Can Googlebot read my website?

Only if two things are true: your robots.txt does not disallow Googlebot, and your server or CDN actually serves the page when Googlebot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Googlebot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is Google-Extended?](https://see-geo.com/bots/google-extended)
- [What is Bingbot?](https://see-geo.com/bots/bingbot)

---

# What is Bingbot?
Source: https://see-geo.com/bots/bingbot · Updated 2026-09-08

Bingbot is Microsoft's search engine crawler. Bing's crawler. Also feeds AI assistants that use Bing's index (including parts of ChatGPT's search), so blocking it hurts more than Bing.

## What does Bingbot do with your content?

Bing's crawler. Also feeds AI assistants that use Bing's index (including parts of ChatGPT's search), so blocking it hurts more than Bing. It follows the rules you publish in robots.txt, a small text file at the root of your site that tells automated visitors what they may read (defined by RFC 9309, the Robots Exclusion Protocol standard).

## Should you block Bingbot?

Blocking Bingbot removes you from more than Bing: AI assistants that draw on Bing's index — including parts of ChatGPT's search experience — lose access to your site too. For almost every business, Bingbot should stay allowed.

## How do you allow or block Bingbot in robots.txt?

Add one of these snippets to the robots.txt file at the root of your site. The first explicitly allows Bingbot everywhere; the second blocks it completely.

```
# Allow Bingbot
User-agent: Bingbot
Allow: /

# Block Bingbot
User-agent: Bingbot
Disallow: /
```

## Is robots.txt enough?

Not always. CDNs and firewalls (Cloudflare in particular) can block crawlers at the network level regardless of what robots.txt says — Cloudflare blocks AI crawlers by default for many accounts. If you want Bingbot to reach your site, check your CDN's bot settings too. SeeGeo's free audit checks both.

## How do I check whether Bingbot can read my site right now?

Two checks, both needed: read your robots.txt for a group naming Bingbot (or the * group it falls back to), then request your homepage with Bingbot's user agent and confirm you get a 200 with your real HTML rather than a 403 or a challenge page. SeeGeo's free crawler check runs both in a few seconds.

[Run the free AI crawler check](https://see-geo.com/ai-crawler-check)

## How does Bingbot compare to similar crawlers?

Bingbot is one of 16 crawlers the SeeGeo audit checks individually. The closest comparisons:

| Crawler | Operator | Role | Runs JavaScript? |
|---|---|---|---|
| Bingbot | Microsoft | search engine crawler | Delayed rendering |
| Googlebot | Google | search engine crawler | Delayed rendering |

[All 16 crawlers, compared](https://see-geo.com/bots) · [What is an AI crawler?](https://see-geo.com/glossary/ai-crawler)

## Why does crawler access matter? The numbers

Crawler access is where AI visibility starts or ends: a blocked crawler can't read you, and an AI that can't read you can't recommend you.

- Visitors arriving from AI assistants convert at roughly 4.4x the rate of traditional organic search on average (Semrush, cross-industry, 2026).
- AI-assistant referrals are still only about 1% of total web traffic — small, but the fastest-growing acquisition channel measured (multiple 2026 studies).
- Adobe Digital Insights (Q1 2026) measured AI-assistant visitors converting 42% better than non-AI traffic — a full reversal from the year before.
- Cloudflare blocks AI crawlers by default for many accounts, so sites are often invisible to AI without anyone having decided to be.
- Content edits alone — statistics, citations, quotable structure — can raise AI visibility on the order of 30–40% (Princeton GEO study, KDD 2024).

## Frequently asked questions

### What is Bingbot?

Bingbot is Microsoft's search engine crawler. Bing's crawler. Also feeds AI assistants that use Bing's index (including parts of ChatGPT's search), so blocking it hurts more than Bing.

### Should I block Bingbot?

Blocking Bingbot removes you from more than Bing: AI assistants that draw on Bing's index — including parts of ChatGPT's search experience — lose access to your site too. For almost every business, Bingbot should stay allowed..

### Does Bingbot run JavaScript?

Bingbot can render JavaScript for indexing, but rendering is delayed and imperfect — server-rendered content is always the safer path.

### Can Bingbot read my website?

Only if two things are true: your robots.txt does not disallow Bingbot, and your server or CDN actually serves the page when Bingbot asks. Firewalls can block it regardless of robots.txt, so the reliable way to know is to fetch your homepage as Bingbot — the free check on see-geo.com/ai-crawler-check does exactly that.

## Related pages

- [What is Googlebot?](https://see-geo.com/bots/googlebot)

---

# AI visibility by industry — practical guides
Source: https://see-geo.com/for · Updated 2026-09-08

How businesses in each industry get recommended when customers ask AI assistants: real prompts, quick wins, and honest expectations.

## Pages

- [AI search visibility for restaurants & cafes](https://see-geo.com/for/restaurants) — How restaurants & cafes get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for bakeries](https://see-geo.com/for/bakeries) — How bakeries get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for SaaS & software companies](https://see-geo.com/for/saas) — How SaaS & software companies get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for agencies & consultancies](https://see-geo.com/for/agencies) — How agencies & consultancies get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for law firms](https://see-geo.com/for/law-firms) — How law firms get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for clinics & healthcare practices](https://see-geo.com/for/healthcare) — How clinics & healthcare practices get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for e-commerce stores](https://see-geo.com/for/ecommerce) — How e-commerce stores get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for contractors & home services](https://see-geo.com/for/contractors) — How contractors & home services get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for gyms & fitness studios](https://see-geo.com/for/fitness) — How gyms & fitness studios get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.
- [AI search visibility for real estate agents](https://see-geo.com/for/real-estate) — How real estate agents get recommended when customers ask AI: the prompts buyers use, industry quick wins, and what to fix first.

---

# AI search visibility for restaurants & cafes
Source: https://see-geo.com/for/restaurants · Updated 2026-09-08

When someone asks an AI assistant to recommend a restaurant, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about restaurants & cafes?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best restaurant for a birthday dinner near me”
- “where should I eat tonight in [city]?”
- “restaurants with good vegan options nearby”
- “is [restaurant name] worth it?”

## Why do AI assistants skip most restaurants & cafes?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for restaurants & cafes

Specific to your industry, roughly in order of effort-to-impact:

- Add Restaurant schema with cuisine, price range, and opening hours — AI assistants lift these fields directly.
- Keep your menu as HTML text, not a PDF or image: crawlers can't read pictures of menus.
- Match your name, address, and phone exactly across your site, Google Business Profile, and Yelp — mismatches make AI unsure you're the same place.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my restaurant recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for restaurants & cafes. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for restaurants & cafes?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

### Do AI assistants use review sites for restaurant recommendations?

Heavily. Answer engines routinely cite Yelp, TripAdvisor, and Google reviews when recommending restaurants, so an unclaimed or thin review profile costs you AI visibility directly.

## Related pages

- [AI search visibility for bakeries](https://see-geo.com/for/bakeries)
- [AI search visibility for SaaS & software companies](https://see-geo.com/for/saas)
- [AI search visibility for agencies & consultancies](https://see-geo.com/for/agencies)

---

# AI search visibility for bakeries
Source: https://see-geo.com/for/bakeries · Updated 2026-09-08

When someone asks an AI assistant to recommend a bakery, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about bakeries?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best custom cake bakery near me”
- “where to order a birthday cake for same-day pickup”
- “wedding cake bakeries in [city] with good reviews”

## Why do AI assistants skip most bakeries?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for bakeries

Specific to your industry, roughly in order of effort-to-impact:

- Publish a page per occasion (weddings, birthdays, corporate) — AI matches specific asks to specific pages, not to a generic gallery.
- State concrete facts AI can quote: years in business, cakes delivered, lead times, price ranges.
- Add Bakery schema (a LocalBusiness subtype) with address, hours, and ordering info.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my bakery recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for bakeries. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for bakeries?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for SaaS & software companies](https://see-geo.com/for/saas)
- [AI search visibility for agencies & consultancies](https://see-geo.com/for/agencies)
- [AI search visibility for law firms](https://see-geo.com/for/law-firms)

---

# AI search visibility for SaaS & software companies
Source: https://see-geo.com/for/saas · Updated 2026-09-08

When someone asks an AI assistant to recommend a SaaS company, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about SaaS & software companies?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best [category] software for a small team”
- “alternatives to [market leader]”
- “[your product] vs [competitor] — which is better?”
- “is [your product] worth it?”

## Why do AI assistants skip most SaaS & software companies?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for SaaS & software companies

Specific to your industry, roughly in order of effort-to-impact:

- Publish honest comparison pages ([you] vs each competitor) with a real feature table — comparison tables are the single most-quoted structure in AI answers about software.
- Maintain a public changelog and pricing page; AI penalizes software it can't price.
- Get listed on G2 and Capterra: answer engines cite both constantly for software queries.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my SaaS company recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for SaaS & software companies. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for SaaS & software companies?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

### Why does my competitor show up in ChatGPT and my product doesn't?

Usually presence, not quality: they're on the comparison sites and 'best X' listicles AI already cites, and you aren't. Getting into the sources AI trusts for your category is the fastest fix.

## Related pages

- [AI search visibility for agencies & consultancies](https://see-geo.com/for/agencies)
- [AI search visibility for law firms](https://see-geo.com/for/law-firms)
- [AI search visibility for clinics & healthcare practices](https://see-geo.com/for/healthcare)

---

# AI search visibility for agencies & consultancies
Source: https://see-geo.com/for/agencies · Updated 2026-09-08

When someone asks an AI assistant to recommend a agency, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about agencies & consultancies?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best marketing agency for startups”
- “should I hire an agency or a freelancer for [task]?”
- “[city] design agencies with e-commerce experience”

## Why do AI assistants skip most agencies & consultancies?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for agencies & consultancies

Specific to your industry, roughly in order of effort-to-impact:

- Publish named case studies with concrete numbers — 'grew organic traffic 3× in 8 months' is liftable; 'we deliver results' is not.
- Niche pages beat a generic services page: one page per industry or problem you serve.
- Put real people on the site: named team members with credentials are an E-E-A-T signal both Google and AI systems weigh.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my agency recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for agencies & consultancies. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for agencies & consultancies?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for law firms](https://see-geo.com/for/law-firms)
- [AI search visibility for clinics & healthcare practices](https://see-geo.com/for/healthcare)
- [AI search visibility for e-commerce stores](https://see-geo.com/for/ecommerce)

---

# AI search visibility for law firms
Source: https://see-geo.com/for/law-firms · Updated 2026-09-08

When someone asks an AI assistant to recommend a law firm, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about law firms?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “do I need a lawyer for [situation]?”
- “best family lawyer near me”
- “how much does a [practice area] lawyer cost?”

## Why do AI assistants skip most law firms?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for law firms

Specific to your industry, roughly in order of effort-to-impact:

- Answer the questions people actually ask ('do I need a lawyer for X?') in plain language — these pages are what AI cites, and they build trust before the call.
- Add Attorney/LegalService schema with practice areas and jurisdiction.
- Publish fee structures, even as ranges: cost pages match high-intent queries AI loves to answer.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my law firm recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for law firms. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for law firms?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

### Is answering legal questions online risky for a law firm?

General educational answers with a clear 'this is not legal advice' note are standard practice and what every top-visibility firm does. Specific advice stays in consultations.

## Related pages

- [AI search visibility for clinics & healthcare practices](https://see-geo.com/for/healthcare)
- [AI search visibility for e-commerce stores](https://see-geo.com/for/ecommerce)
- [AI search visibility for contractors & home services](https://see-geo.com/for/contractors)

---

# AI search visibility for clinics & healthcare practices
Source: https://see-geo.com/for/healthcare · Updated 2026-09-08

When someone asks an AI assistant to recommend a clinic, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about clinics & healthcare practices?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “dentist near me that takes [insurance]”
- “how much does [procedure] cost without insurance?”
- “best physical therapist for back pain in [city]”

## Why do AI assistants skip most clinics & healthcare practices?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for clinics & healthcare practices

Specific to your industry, roughly in order of effort-to-impact:

- List accepted insurance plans as HTML text — it's one of the most-asked questions and most sites bury it in a PDF.
- Add Physician/MedicalClinic schema with specialties and hours.
- Publish honest cost-range pages for common procedures; they match the highest-intent queries in healthcare.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my clinic recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for clinics & healthcare practices. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for clinics & healthcare practices?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for e-commerce stores](https://see-geo.com/for/ecommerce)
- [AI search visibility for contractors & home services](https://see-geo.com/for/contractors)
- [AI search visibility for gyms & fitness studios](https://see-geo.com/for/fitness)

---

# AI search visibility for e-commerce stores
Source: https://see-geo.com/for/ecommerce · Updated 2026-09-08

When someone asks an AI assistant to recommend a online store, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about e-commerce stores?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best [product] under $[price]”
- “[product A] vs [product B] — which should I buy?”
- “is [store] legit?”

## Why do AI assistants skip most e-commerce stores?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for e-commerce stores

Specific to your industry, roughly in order of effort-to-impact:

- Product schema with price, availability, and aggregate ratings is non-negotiable — AI shopping answers are assembled from it.
- Publish genuinely useful buying guides for your category; guides get cited where product pages don't.
- Keep an honest, findable shipping & returns page: 'is [store] legit' answers are built from trust signals like it.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my online store recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for e-commerce stores. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for e-commerce stores?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for contractors & home services](https://see-geo.com/for/contractors)
- [AI search visibility for gyms & fitness studios](https://see-geo.com/for/fitness)
- [AI search visibility for real estate agents](https://see-geo.com/for/real-estate)

---

# AI search visibility for contractors & home services
Source: https://see-geo.com/for/contractors · Updated 2026-09-08

When someone asks an AI assistant to recommend a contractor, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about contractors & home services?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “roof repair cost for a [size] house”
- “best plumber near me for an emergency”
- “how do I know if a contractor quote is fair?”

## Why do AI assistants skip most contractors & home services?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for contractors & home services

Specific to your industry, roughly in order of effort-to-impact:

- Publish cost-guide pages with real ranges for your services — cost questions dominate home-services AI queries.
- List your license numbers and insurance visibly; AI treats verifiable credentials as trust signals.
- One page per service area (city/suburb) with genuinely local detail — thin duplicated location pages get ignored.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my contractor recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for contractors & home services. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for contractors & home services?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for gyms & fitness studios](https://see-geo.com/for/fitness)
- [AI search visibility for real estate agents](https://see-geo.com/for/real-estate)
- [AI search visibility for restaurants & cafes](https://see-geo.com/for/restaurants)

---

# AI search visibility for gyms & fitness studios
Source: https://see-geo.com/for/fitness · Updated 2026-09-08

When someone asks an AI assistant to recommend a gym, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about gyms & fitness studios?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best gym near me with childcare”
- “how much does [gym] cost per month?”
- “pilates vs yoga for beginners — where should I start?”

## Why do AI assistants skip most gyms & fitness studios?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for gyms & fitness studios

Specific to your industry, roughly in order of effort-to-impact:

- Publish your actual prices — 'how much does it cost' is the top question, and hiding pricing hands the AI answer to a competitor who publishes theirs.
- Add HealthClub schema with amenities, classes, and hours.
- Answer beginner questions ('pilates vs yoga') on your site; educational pages earn the citation and the visit.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my gym recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for gyms & fitness studios. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for gyms & fitness studios?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for real estate agents](https://see-geo.com/for/real-estate)
- [AI search visibility for restaurants & cafes](https://see-geo.com/for/restaurants)
- [AI search visibility for bakeries](https://see-geo.com/for/bakeries)

---

# AI search visibility for real estate agents
Source: https://see-geo.com/for/real-estate · Updated 2026-09-08

When someone asks an AI assistant to recommend a real estate agency, the AI names two or three businesses and most owners have no idea whether they're one of them. Being named is not luck — it's the direct result of how readable, structured, and well-referenced your business is, and all three are fixable.

## What do customers ask AI about real estate agents?

These are the shapes of question your future customers put to ChatGPT, Claude, Gemini, and Google's AI Overviews. Every one of them ends with the AI recommending someone — the only question is whether it's you:

- “best real estate agent in [neighborhood]”
- “what's my home worth in [city]?”
- “should I sell my house now or wait?”

## Why do AI assistants skip most real estate agents?

Three reasons, in order of frequency. First, the AI can't read the site: most AI crawlers don't run JavaScript and some are blocked by default at the CDN level, so the site is effectively blank to them. Second, the site never plainly says what the business is and where it operates, so the AI can't confidently classify it. Third, the business is absent from the sources the AI already cites for these questions — directories, review platforms, and listicles.

Research on AI visibility (the Princeton GEO study, presented at KDD 2024) suggests the highest-leverage content edits are adding concrete statistics and citing sources — improvements on the order of 30–40% in how often content is used in AI answers. Not a guarantee, but the best-evidenced starting point available.

## Quick wins for real estate agents

Specific to your industry, roughly in order of effort-to-impact:

- Publish neighborhood guides with concrete data (median prices, days on market) — the pages AI cites for local market questions.
- Add RealEstateAgent schema with service areas.
- Named-agent pages with credentials, sales history, and reviews beat an anonymous team page for both trust and AI attribution.

## How do you know if it's working?

Measure, don't guess: AI answers don't appear in your analytics, so the only way to know whether you're being recommended is to ask the engines the same questions your customers do, on a schedule, and track the trend. That's what SeeGeo's visibility scans do across ChatGPT, Claude, Gemini, and Google — and a single scan is a snapshot, never a score; trends over weeks are what matter.

## Frequently asked questions

### How do I get my real estate agency recommended by ChatGPT?

Make your site readable to AI crawlers (no JavaScript-only content, no accidental blocks), state plainly what you do and where, add structured data, and get present on the review sites and directories AI already cites for real estate agents. Then measure whether it's working by asking the engines your customers' questions on a schedule.

### Does normal SEO still matter for real estate agents?

Yes — search and AI answers overlap heavily, and Google's AI Overviews are built on the same crawl as Google Search. The work compounds: most fixes that help you rank also help you get cited.

### How is GEO different from SEO?

SEO optimizes for ranking in a list of links; GEO (generative engine optimization) optimizes for being named and cited inside an AI-written answer. GEO is younger and less settled than SEO — anyone promising guaranteed AI rankings is overselling.

## Related pages

- [AI search visibility for restaurants & cafes](https://see-geo.com/for/restaurants)
- [AI search visibility for bakeries](https://see-geo.com/for/bakeries)
- [AI search visibility for SaaS & software companies](https://see-geo.com/for/saas)

---

# The GEO glossary: AI visibility, defined term by term
Source: https://see-geo.com/glossary · Updated 2026-09-08

Every term you'll meet in AI visibility work — GEO, grounding, AI citations, share of voice — each with a liftable definition and the 'so what'.

## Pages

- [What is Query fan-out?](https://see-geo.com/glossary/query-fan-out) — Query fan-out: how AI search splits one question into sub-queries and merges the sources. Which pages get cited, and how to write for it.
- [What is llms.txt?](https://see-geo.com/glossary/llms-txt) — llms.txt explained: what the file is, who proposed it, whether ChatGPT, Claude or Google read it, and whether it is worth adding to your site in 2026.
- [What is AI visibility audit?](https://see-geo.com/glossary/ai-visibility-audit) — AI visibility audit, defined: the six things it measures, how it differs from an SEO audit, what a score means and how to run one free in about 20 seconds.
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization) — GEO means optimizing your site so AI assistants can read, quote, and recommend it. How it differs from SEO, and what actually moves it.
- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility) — AI visibility is how often AI assistants mention or recommend you. What it's made of, why it differs from rankings, and how to measure it honestly.
- [What is Answer engine?](https://see-geo.com/glossary/answer-engine) — An answer engine replies with a synthesized answer, not links. Why that changes visibility, and what it reads to build its answers.
- [What is AI crawler?](https://see-geo.com/glossary/ai-crawler) — AI crawlers read your site for AI training, AI search, or live answers. The three kinds, why most don't run JavaScript, and how access decides visibility.
- [What is Google AI Overviews?](https://see-geo.com/glossary/ai-overviews) — Google's AI Overviews sit above search results and run on Googlebot's index. What decides inclusion, and what blocking Google-Extended does and doesn't do.
- [What is Grounding?](https://see-geo.com/glossary/grounding) — Grounding is an AI answering from live retrieved pages instead of memory. Why it's rarer than assumed, and why that decides how citations work.
- [What is AI citation?](https://see-geo.com/glossary/ai-citation) — An AI citation is a source credit inside a generated answer. Why cited pages are mostly third-party, and what earns a citation for your own pages.
- [What is Share of voice (AI)?](https://see-geo.com/glossary/share-of-voice) — AI share of voice: how often answers name you vs competitors across repeated sampling. Why it beats any single 'AI rank', and how to measure it.
- [What is Prompt panel?](https://see-geo.com/glossary/prompt-panel) — A prompt panel is the fixed question set behind honest AI-visibility tracking. What a good panel contains and why it must stay stable.
- [What is Content extractability?](https://see-geo.com/glossary/content-extractability) — Extractability is whether AI can lift a self-contained answer from your page. The measured levers: answer-first leads, question headings, stats, dates.
- [What is Entity clarity?](https://see-geo.com/glossary/entity-clarity) — Entity clarity: whether machines can tell what your business is. Why visible prose must match your structured data, and which pages anchor identity.
- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data) — Structured data (JSON-LD) tells machines what your pages are. Which types matter for AI visibility, and the one rule: never mark up what isn't visible.
- [What is robots.txt?](https://see-geo.com/glossary/robots-txt) — robots.txt controls which crawlers may read your site — now including every AI crawler. The syntax, the accidents, and what it can't control.
- [What is LLM training data?](https://see-geo.com/glossary/llm-training-data) — Training data decides what AI models know about you from memory. How the web gets in, the years-long lag, and the block-or-allow trade-off.
- [What is E-E-A-T?](https://see-geo.com/glossary/e-e-a-t) — E-E-A-T is the trust checklist search raters use — and AI answers increasingly mirror. The signals that are checkable by machine, and the ones that aren't.
- [What is Answer-first content?](https://see-geo.com/glossary/answer-first-content) — Answer-first content puts the full answer in the first two paragraphs. Why inverted structure wins AI citations, and how to retrofit existing pages.

---

# What is Query fan-out?
Source: https://see-geo.com/glossary/query-fan-out · Updated 2026-09-08

Query fan-out is how AI search systems answer one question: they silently expand it into several related sub-queries, retrieve sources for each, and assemble a single answer from the results — so a page can be cited for a question it never literally contains.

## How does query fan-out work in AI search?

When you ask Google's AI Mode or ChatGPT search a question like 'best accountant for a small café in Lyon', the system does not run that string once. It generates a fan of narrower queries — accountants in Lyon, accounting for hospitality, small-business bookkeeping fees — runs each against the web index, and synthesises one answer from the union of sources. Google described the mechanism publicly when it launched AI Mode in 2025; Perplexity and ChatGPT search behave the same way.

The consequence is that citation is decided per sub-query. A page that answers one narrow facet very plainly can be cited in the final answer even though it never mentions the original question.

## What does query fan-out change about how to write?

It rewards pages that answer one specific question completely, under a heading phrased as that question, with a self-contained paragraph an engine can lift. A long page that covers ten things vaguely loses to ten short sections that each own one sub-query. It also rewards covering the adjacent questions a customer would ask next — those are the sub-queries the fan-out produces.

[Our full explainer on query fan-out](https://see-geo.com/blog/query-fan-out-explained)

## How do you know which sub-queries you are cited for?

You cannot see the fan-out directly, but you can measure its output: ask the assistants the questions your customers ask and record whether you are named and which pages are cited. SeeGeo's tracking does this on a schedule across ChatGPT, Claude and Gemini, and the audit's extractability score tells you whether a page has the shape a sub-query can lift.

## Frequently asked questions

### What does Query fan-out mean?

Query fan-out is how AI search systems answer one question: they silently expand it into several related sub-queries, retrieve sources for each, and assemble a single answer from the results — so a page can be cited for a question it never literally contains.

### What is query fan-out in AI search?

Query fan-out is the step where an AI search engine turns one question into several related sub-queries, retrieves sources for each, and merges them into one answer. It is why a page can be cited for a question it never states word for word.

### Does query fan-out mean I should write longer pages?

No — it means writing more precisely. Each sub-query is matched to a passage, so pages that answer one question plainly under a question-shaped heading are lifted more often than long pages that mention many things.

### Which engines use query fan-out?

Google AI Mode and AI Overviews describe it explicitly; ChatGPT search, Perplexity and Claude with web search retrieve in the same multi-query way even where they do not use the term.

## Related pages

- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is AI citation?](https://see-geo.com/glossary/ai-citation)
- [What is Answer-first content?](https://see-geo.com/glossary/answer-first-content)

---

# What is llms.txt?
Source: https://see-geo.com/glossary/llms-txt · Updated 2026-09-08

llms.txt is a proposed plain-text file at the root of a website that summarises the site for large language models and points them to its most useful pages — a suggestion to AI systems, not an instruction, and as of 2026 not confirmed to be read by any major assistant.

## Is llms.txt worth adding to my website?

Honest answer: it costs almost nothing and helps almost nothing measurable yet. No major assistant — OpenAI, Anthropic, Google, Perplexity — has confirmed reading llms.txt, and our own study of real small-business sites found no visibility difference between sites with and without one. Add it if you have ten spare minutes; do not expect it to move a score.

What does move a score is the content the file would point to: pages that open with an answer, headings phrased as questions, structured data that says what you are, and crawlers that are actually allowed in. Those are what assistants read today.

[Our full verdict on llms.txt](https://see-geo.com/blog/llms-txt-honest-verdict) · [Generate an llms.txt in a minute](https://see-geo.com/llms-txt-generator)

## How is llms.txt different from robots.txt?

robots.txt is a standard every major crawler obeys; it says where bots may not go. llms.txt is a proposal from 2024 that says where models might like to look. One is enforced by the crawlers themselves, the other depends entirely on a reader choosing to honour it. Blocking a bot in robots.txt while inviting it in llms.txt achieves nothing.

## What should an llms.txt contain if you add one?

A one-paragraph description of the business in plain words, then a short list of your most useful pages with one-line descriptions each — the About page, the pricing page, the FAQ, the two or three guides customers actually need. Keep it under a screen; it is a map, not a copy of the site.

## Frequently asked questions

### What does llms.txt mean?

llms.txt is a proposed plain-text file at the root of a website that summarises the site for large language models and points them to its most useful pages — a suggestion to AI systems, not an instruction, and as of 2026 not confirmed to be read by any major assistant.

### Is llms.txt worth adding to my website?

It is harmless and cheap, but as of 2026 no major AI assistant confirms reading it and we measured no visibility difference for sites that have one. Spend the time on crawler access, answer-first content and structured data first.

### Do ChatGPT, Claude or Google read llms.txt?

None of them has said so. Anthropic publishes its own documentation in the format, and some developer tools read it, but there is no evidence the consumer assistants consult it when answering a question about your business.

### Can llms.txt replace robots.txt?

No. robots.txt is the enforced standard for what crawlers may fetch; llms.txt is a voluntary summary. They answer different questions and only one of them is obeyed.

## Related pages

- [What is robots.txt?](https://see-geo.com/glossary/robots-txt)
- [What is AI crawler?](https://see-geo.com/glossary/ai-crawler)
- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data)

---

# What is AI visibility audit?
Source: https://see-geo.com/glossary/ai-visibility-audit · Updated 2026-09-08

An AI visibility audit is a check of whether AI assistants — ChatGPT, Claude, Gemini, Perplexity, Google's AI features — can find, read, understand and cite a website, scored across crawler access, technical foundation, structured data, content extractability, off-site presence and entity clarity.

## What does an AI visibility audit measure?

Six categories, each answering a question a machine asks about your site. Access: can AI crawlers reach the pages, given robots.txt and any firewall? Technical foundation: is the text there without JavaScript, are the URLs clean, does the page load? Structured data: does the site declare what it is in a form machines cannot misread? Content extractability: does each page open with an answer, use question-shaped headings, carry a date and sources? Off-site presence: is the business in Wikipedia, Wikidata and the places assistants cite? Entity clarity: can a classifier tell what kind of business this is and where?

Each category is scored 0–100 and combined into one grade. In SeeGeo's engine the grade is anchored to the score distribution of 107 real small-business websites, so an A is the top decile of real sites.

[Run the free audit](https://see-geo.com/geo-audit)

## How is an AI visibility audit different from an SEO audit?

An SEO audit asks whether Google can rank your pages; an AI visibility audit asks whether an assistant would read and cite them when answering a customer. The overlap is real — crawlable, fast, well-structured pages help both — but AI audits weigh things SEO tools ignore: per-bot crawler permissions, firewalls that block AI user agents, whether text survives without JavaScript, and whether a paragraph is liftable as an answer.

## What does the audit not measure?

Whether assistants actually name you today. That is a separate, ongoing measurement — asking the assistants your customers' questions on a schedule and recording mentions and citations — because the answer changes week to week and depends on competitors, not just on your site. The audit measures readiness; tracking measures outcome.

[What AI visibility means](https://see-geo.com/glossary/ai-visibility)

## Frequently asked questions

### What does AI visibility audit mean?

An AI visibility audit is a check of whether AI assistants — ChatGPT, Claude, Gemini, Perplexity, Google's AI features — can find, read, understand and cite a website, scored across crawler access, technical foundation, structured data, content extractability, off-site presence and entity clarity.

### What does an AI visibility audit measure?

Whether AI assistants can find, read, understand and cite your site: crawler access, technical foundation, structured data, content extractability, off-site presence and entity clarity, each scored 0–100 and combined into a grade.

### How long does an AI visibility audit take?

SeeGeo's audit crawls the homepage and up to about eight pages the way an AI crawler does and returns the graded report in roughly 20 seconds, free and without an account for the score.

### Is an AI visibility audit the same as GEO audit?

Yes — GEO (generative engine optimization) audit and AI visibility audit describe the same check. Some tools also call it an AEO or answer-engine audit.

## Related pages

- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is Content extractability?](https://see-geo.com/glossary/content-extractability)

---

# What is Generative Engine Optimization (GEO)?
Source: https://see-geo.com/glossary/generative-engine-optimization · Updated 2026-09-08

Generative Engine Optimization (GEO) is the practice of making a website visible, quotable, and citable to AI systems that generate answers — ChatGPT, Claude, Gemini, Google AI Overviews — the way SEO makes it visible to search engines.

## How is GEO different from SEO?

SEO earns a ranked position on a results page; GEO earns a place inside a generated answer. The mechanics overlap — both need crawlable, understandable content — but the goals diverge: a search engine ranks pages, while an answer engine assembles a response from sources and may name only two or three. There is no position eleven in an AI answer: you are in it or you are absent.

The disciplines also diverge on what content wins. Search rewards depth and links; generated answers reward liftable statements — definitions, statistics with sources, direct answers under question-shaped headings — because those are the shapes a model can quote without rewriting.

## What actually improves GEO?

The measured levers, in rough order: let AI crawlers in (many sites block them by accident at the CDN), serve real content without JavaScript (most AI crawlers don't run it), state plainly what you are in visible prose, structure content so single passages answer single questions, and be present on the third-party pages AI already cites — because answers are assembled mostly from those, not from vendors' own sites.

- Content edits alone — statistics, citations, quotable structure — raised AI visibility on the order of 30–40% in the Princeton GEO study (KDD 2024).
- Visitors from AI assistants convert at roughly 4.4x organic search on average (Semrush, 2026).

[Run a free GEO audit](https://see-geo.com/) · [SEO vs GEO, in depth](https://see-geo.com/blog/seo-vs-geo-difference)

## What's the most common GEO mistake?

Treating GEO as wording tweaks while the plumbing is broken. Rewriting copy on a site whose CDN blocks AI crawlers, or whose content only renders with JavaScript, optimizes pages no engine will ever read. The order of operations is fixed: access first, extractability second, off-site presence third — and an audit tells you in seconds which stage you're actually at.

## Frequently asked questions

### What does Generative Engine Optimization (GEO) mean?

Generative Engine Optimization (GEO) is the practice of making a website visible, quotable, and citable to AI systems that generate answers — ChatGPT, Claude, Gemini, Google AI Overviews — the way SEO makes it visible to search engines.

### Is GEO replacing SEO?

No — they compound. AI answers are built partly on search infrastructure (AI Overviews run on Googlebot's index, ChatGPT search draws on Bing), so classic SEO remains the foundation GEO builds on. The sites winning AI visibility tend to be doing both.

### Is GEO the same as AEO (Answer Engine Optimization)?

Effectively yes — AEO, GEO, and 'LLM SEO' all name the same practice: optimizing content to be read, quoted, and recommended by AI systems. GEO is the term the research literature settled on.

### Can GEO be measured?

Yes, by sampling: ask the engines a fixed panel of questions repeatedly and record whether you're mentioned, where you rank in ranked answers, and whether your pages are cited. Because answers vary between runs, honest measurement reports rates across repeated samples, never a single 'AI rank'.

## Related pages

- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility)
- [What is Answer engine?](https://see-geo.com/glossary/answer-engine)
- [What is Content extractability?](https://see-geo.com/glossary/content-extractability)

---

# What is AI visibility?
Source: https://see-geo.com/glossary/ai-visibility · Updated 2026-09-08

AI visibility is how often and how favorably a business appears in AI-generated answers when people ask assistants like ChatGPT, Claude, or Gemini for information or recommendations in its category.

## What is AI visibility made of?

Three measurable layers. Mention: does the answer name you at all? Position: when the answer is a ranked list, where do you sit? Citation: does the assistant link or credit your pages as a source? A brand can score well on one and badly on another — being named in every answer while your website is never actually read is common, because answers are assembled largely from third-party pages like comparison posts and directories.

## Why can't AI visibility be a single rank?

Because generated answers are non-deterministic: the same question asked twice can produce differently ordered, differently worded answers. Any tool selling a fixed 'AI rank' is measuring noise. Honest measurement samples the same questions repeatedly and reports rates — mentioned in 58 of 60 runs — the way polling reports margins rather than certainties.

[How SeeGeo measures it](https://see-geo.com/about) · [Does AI visibility drive revenue?](https://see-geo.com/blog/ai-visibility-revenue-data-2026)

## What's the most common measurement mistake?

Measuring once. One prompt asked one time produces an anecdote — thrilling or alarming, and equally meaningless, because the same question re-asked can flip the answer. The second most common: measuring only branded prompts, which tests whether engines know your name rather than whether they recommend you to strangers. Fixed panels, repeated runs, mostly unbranded questions.

## Frequently asked questions

### What does AI visibility mean?

AI visibility is how often and how favorably a business appears in AI-generated answers when people ask assistants like ChatGPT, Claude, or Gemini for information or recommendations in its category.

### How do I check my AI visibility for free?

Ask the assistants what your customers would ask — category questions, not your brand name — and note whether you appear. For the technical half, run a crawlability and extractability audit like SeeGeo's free one: if AI crawlers can't read your site, visibility can't follow.

### How much traffic does AI visibility actually drive?

Around 1% of total web traffic today across industries — small, but the fastest-growing channel measured, and the visitors convert at roughly 4.4x organic search on average (Semrush, 2026). The volume argument is about the trend, not the present.

### Why am I visible in Google but not in AI answers?

The usual causes: your CDN blocks AI crawlers even though Googlebot passes, your content only renders with JavaScript (most AI crawlers don't run it), or the third-party pages AI builds answers from — comparisons, directories — don't mention you even though your own SEO is fine.

## Related pages

- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is Share of voice (AI)?](https://see-geo.com/glossary/share-of-voice)
- [What is AI citation?](https://see-geo.com/glossary/ai-citation)

---

# What is Answer engine?
Source: https://see-geo.com/glossary/answer-engine · Updated 2026-09-08

An answer engine is a system that responds to a question with a synthesized answer instead of a list of links — ChatGPT, Claude, Perplexity, and Google's AI Overviews are answer engines; classic Google Search is not.

## How does an answer engine choose its sources?

Two routes. From memory: the model answers from what it learned in training, which favors brands that were well-documented on the web when the model was trained. From retrieval: the engine runs a live search, reads a handful of pages, and assembles the answer from them — favoring pages that are crawlable, fast, and structured so a passage can be lifted whole.

In practice retrieval reads surprisingly few pages — often under ten per answer — and they are disproportionately third-party: comparison posts, directories, community threads. Being absent from those pages means being absent from the answer, however good your own site is.

[What is grounding?](https://see-geo.com/glossary/grounding)

## Which answer engines matter for a business?

By reach today: Google AI Overviews (shown above classic results to Google's billions of users), ChatGPT, Gemini, Claude, and Perplexity. They share an important property: all of them read the web through crawlers you can allow or block, which makes crawler access the first, cheapest thing to get right.

| Engine | Operator | Reads the web via |
|---|---|---|
| AI Overviews | Google | Googlebot |
| ChatGPT | OpenAI | OAI-SearchBot / GPTBot / Bing's index |
| Gemini | Google | Googlebot / Google-Extended |
| Claude | Anthropic | ClaudeBot / Claude-User |
| Perplexity | Perplexity | PerplexityBot / Perplexity-User |

[Every crawler, explained](https://see-geo.com/bots)

## What's the most common answer-engine mistake?

Optimizing only your own website. Answer engines assemble responses largely from third-party pages — comparisons, directories, community threads — so a perfect site that's absent from those sources still loses the answer. Your site earns the citation when retrieval reads it; the third-party web decides whether you're in the conversation at all. Budget effort on both.

## Frequently asked questions

### What does Answer engine mean?

An answer engine is a system that responds to a question with a synthesized answer instead of a list of links — ChatGPT, Claude, Perplexity, and Google's AI Overviews are answer engines; classic Google Search is not.

### Do answer engines send traffic?

Less volume than search, better visitors: studies through 2026 consistently find AI-referred visitors convert several times better than organic — Semrush's cross-industry figure is about 4.4x — because they arrive pre-qualified by the answer that sent them.

### Can I be in AI answers without being in Google?

Rarely. Most answer engines lean on search indexes for retrieval — AI Overviews on Google's, parts of ChatGPT on Bing's — so classic indexing remains the foundation. The reverse is common though: ranking in Google while absent from answers.

### Is Perplexity an answer engine?

Yes, and the most retrieval-heavy of the majors: nearly every Perplexity answer cites the specific pages it read, which makes it the clearest window into which sources answer engines actually trust for your category.

## Related pages

- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is Grounding?](https://see-geo.com/glossary/grounding)
- [What is Google AI Overviews?](https://see-geo.com/glossary/ai-overviews)

---

# What is AI crawler?
Source: https://see-geo.com/glossary/ai-crawler · Updated 2026-09-08

An AI crawler is an automated program that reads web pages on behalf of an AI system — for training data (GPTBot, ClaudeBot), for an AI search index (OAI-SearchBot, PerplexityBot), or live, mid-conversation (ChatGPT-User, Claude-User).

## What are the three kinds of AI crawler?

Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) collect text that future models learn from — blocking them trades future model knowledge of your business for keeping content out of training data. Index crawlers (OAI-SearchBot, PerplexityBot) feed AI search products; blocking them removes you from those answers directly. Live fetchers (ChatGPT-User, Claude-User, Perplexity-User) retrieve a page on demand when an assistant needs it mid-answer.

[All 16 crawlers, individually explained](https://see-geo.com/bots)

## Why does JavaScript rendering matter so much?

Nearly all AI crawlers read raw HTML and do not execute JavaScript. A site that paints its content client-side — common with React, Vue, and site builders — serves those crawlers an effectively blank page. This is the single most damaging finding SeeGeo's audit produces, because it silently zeroes out every other effort: to an AI, the site does not say anything at all.

## How do sites block AI crawlers by accident?

Two ways: leftover robots.txt rules written for another purpose, and CDN bot protection — Cloudflare in particular blocks AI crawlers by default for many accounts, at the network level, where robots.txt can't help. Sites are routinely invisible to AI without anyone having decided to be. An audit that checks each crawler individually is how you find out.

[Check your site free](https://see-geo.com/)

## Frequently asked questions

### What does AI crawler mean?

An AI crawler is an automated program that reads web pages on behalf of an AI system — for training data (GPTBot, ClaudeBot), for an AI search index (OAI-SearchBot, PerplexityBot), or live, mid-conversation (ChatGPT-User, Claude-User).

### Should I block AI crawlers?

It's a real trade-off, not a default. Blocking training crawlers keeps content out of future models at the cost of those models knowing you exist; blocking index crawlers and live fetchers removes you from AI answers outright. If being recommended matters to your business, most crawlers should stay allowed.

### Do AI crawlers respect robots.txt?

The major ones — from OpenAI, Anthropic, Google, Perplexity — publicly commit to honoring robots.txt (RFC 9309) and in practice do. The practical problem is usually the opposite: sites blocking crawlers unintentionally, not crawlers ignoring rules.

### How do I see which AI crawlers can read my site?

Check robots.txt for each crawler's user agent, then check what your CDN does at the network layer — the part robots.txt can't show. SeeGeo's free audit tests all sixteen major search and AI crawlers individually and reports each one's access.

## Related pages

- [What is robots.txt?](https://see-geo.com/glossary/robots-txt)
- [What is LLM training data?](https://see-geo.com/glossary/llm-training-data)
- [What is Content extractability?](https://see-geo.com/glossary/content-extractability)

---

# What is Google AI Overviews?
Source: https://see-geo.com/glossary/ai-overviews · Updated 2026-09-08

AI Overviews are the AI-generated summaries Google shows above classic search results, built on Googlebot's index — meaning your access to them is decided by ordinary Google crawling, not by any separate AI crawler.

## How do you appear in AI Overviews?

The prerequisites are classic SEO: indexed by Googlebot, served without JavaScript dependence, structured so a passage can answer a question on its own. Overviews cite sources, and studies through 2026 find the cited pages skew toward content with liftable structure — direct answers, statistics, clear headings — over pages that merely rank well.

## Does blocking Google-Extended remove you from AI Overviews?

No — this is the most common confusion in the category. Google-Extended controls Gemini model training only. AI Overviews run on Googlebot, so the only way to opt out of them is to opt out of Google Search itself. You can block Gemini training and keep full Overview visibility.

[Google-Extended, explained](https://see-geo.com/bots/google-extended) · [Googlebot, explained](https://see-geo.com/bots/googlebot)

## Do AI Overviews appear for every search?

No — they appear for a subset of queries, skewed toward informational and comparison questions, and their frequency varies by market, language, and Google's ongoing experiments. That volatility is worth internalizing: a visibility strategy built only on Overviews inherits Google's experiment schedule, while the underlying fundamentals — indexability, extractability, entity clarity — pay off across every answer engine at once.

## Frequently asked questions

### What does Google AI Overviews mean?

AI Overviews are the AI-generated summaries Google shows above classic search results, built on Googlebot's index — meaning your access to them is decided by ordinary Google crawling, not by any separate AI crawler.

### Do AI Overviews reduce website traffic?

For informational queries they can — the answer is on the results page. But cited sources gain a new click path, and the visitors who do click through convert better; the strategy is to be the cited source rather than the summarized one.

### Can I opt out of AI Overviews but stay in Google Search?

Not cleanly. Overviews are built on the same Googlebot index as search, so there's no separate crawler to block. Preventing snippet reuse via nosnippet directives limits how your content appears, at a real cost to how search presents you.

### Which crawler controls Gemini, and which controls Overviews?

Gemini training is controlled by Google-Extended; AI Overviews and Google Search both run on Googlebot. Blocking Google-Extended affects only Gemini training — rankings and Overviews are untouched.

## Related pages

- [What is Answer engine?](https://see-geo.com/glossary/answer-engine)
- [What is AI crawler?](https://see-geo.com/glossary/ai-crawler)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)

---

# What is Grounding?
Source: https://see-geo.com/glossary/grounding · Updated 2026-09-08

Grounding is when an AI model backs its answer with live retrieved sources — running a web search mid-answer and reading real pages — instead of answering purely from what it memorized in training.

## How often do AI answers actually ground?

Far less than assumed. Models ground when they judge a question to need current information; for evergreen category questions — 'best X for Y' — they frequently answer from memory. In SeeGeo's own measurement, a model offered live search used it on 1 of 20 category prompts, answering the other 19 from training data alone. An ungrounded answer carries no citations, so 'not cited' often really means 'nothing was looked up'.

## Why does grounding matter for visibility?

Because it splits visibility into two games. Ungrounded answers draw on training data — you win those by being well-documented on the web over time, which is slow. Grounded answers draw on a handful of freshly retrieved pages — you win those by being present on the exact pages retrieval selects, mostly third-party comparisons and directories. A serious visibility strategy plays both.

[What is an AI citation?](https://see-geo.com/glossary/ai-citation)

## Can you make an answer ground?

Sometimes — questions that demand current information ('pricing in 2026', 'latest') trigger search far more often than evergreen ones. But you don't control how customers phrase things, so treating grounding as a switch you can flip is a mistake. The robust posture plays both games: be well-documented enough to win ungrounded answers from memory, and present on retrieval-trusted pages for the grounded ones.

## Frequently asked questions

### What does Grounding mean?

Grounding is when an AI model backs its answer with live retrieved sources — running a web search mid-answer and reading real pages — instead of answering purely from what it memorized in training.

### How can I tell if an answer was grounded?

Citations are the tell: grounded answers can link the pages they read, ungrounded ones have nothing to link. Perplexity grounds nearly always; ChatGPT and Gemini decide per question, which is why the same prompt sometimes returns sources and sometimes none.

### Does grounding use my robots.txt rules?

Yes — grounded retrieval reaches your site through the engine's crawlers and fetchers (OAI-SearchBot, ChatGPT-User, Google's infrastructure), all of which honor crawler rules. Blocking them blocks grounded answers from ever reading you.

### Why did my citation disappear between two identical questions?

Most likely the first answer grounded and the second didn't — the model chose memory the second time, and an answer from memory has no citations to give. This is why honest citation tracking distinguishes 'not cited' from 'nothing was retrieved at all'.

## Related pages

- [What is AI citation?](https://see-geo.com/glossary/ai-citation)
- [What is Answer engine?](https://see-geo.com/glossary/answer-engine)
- [What is LLM training data?](https://see-geo.com/glossary/llm-training-data)

---

# What is AI citation?
Source: https://see-geo.com/glossary/ai-citation · Updated 2026-09-08

An AI citation is a source link or credit inside a generated answer, pointing at a page the engine actually retrieved and used — the AI-era equivalent of ranking on page one.

## Which pages do AI answers actually cite?

Disproportionately third-party ones. When SeeGeo probed the sources behind category answers for a market-leading product, the leader's own website appeared in zero of the retrieved pages — while comparison listicles, YouTube, Reddit, and competitors' comparison pages were read repeatedly. Being named in the answer and being read for the answer are different achievements: the first follows fame, the second follows being on the pages retrieval trusts.

[The full case study](https://see-geo.com/blog/we-audited-ourselves)

## What makes a page citable?

Retrieval-reachable first: crawlable, fast, readable without JavaScript. Then liftable: a citable page answers a specific question in a self-contained passage — definitions, statistics with named sources, direct comparisons. Engines cite what they can quote cleanly; pages that require assembling meaning across paragraphs get read and then not credited.

- Adding statistics and quotable structure raised visibility 30–40% in the Princeton GEO study (KDD 2024).
- Glossary and definition pages are among the most-cited content shapes — a definition is the most liftable passage there is.

## What's the most common citation mistake?

Reading 'no citations' as 'not cited'. An answer that never grounded has no citations to give anyone — counting it as a loss corrupts the metric and sends you fixing the wrong thing. The second mistake: pouring effort into making your own pages citable while ignoring that most retrieved pages are third-party. Both errors come from skipping the same question: did the engine retrieve anything, and if so, what?

## Frequently asked questions

### What does AI citation mean?

An AI citation is a source link or credit inside a generated answer, pointing at a page the engine actually retrieved and used — the AI-era equivalent of ranking on page one.

### Why is my brand mentioned but never cited?

Mentions come from training data — the model knows you exist. Citations come from retrieval — your pages must be among the few an engine reads live for that answer. The gap usually means the third-party pages retrieval prefers don't feature you, even though your reputation reached the training data.

### Do citations in AI answers drive clicks?

Fewer than a top ranking, better than none — and the citation itself carries authority: being the named source in an answer positions you as the reference, which compounds as more answers are generated from the same retrieval preferences.

### How do I track AI citations honestly?

Sample fixed questions repeatedly and record which pages the answers credit, separating 'not cited' from 'the answer never retrieved anything' — an ungrounded answer can't cite anyone, and counting it as a citation loss corrupts the metric.

## Related pages

- [What is Grounding?](https://see-geo.com/glossary/grounding)
- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility)
- [What is Share of voice (AI)?](https://see-geo.com/glossary/share-of-voice)

---

# What is Share of voice (AI)?
Source: https://see-geo.com/glossary/share-of-voice · Updated 2026-09-08

In AI visibility, share of voice is how often you are named in AI answers for your category's questions relative to your competitors — the answer-engine equivalent of market share of attention.

## How is AI share of voice measured?

Fix a panel of the questions your customers actually ask, run each against the engines several times (answers vary between runs), and count appearances for you and each competitor. The output reads like polling: named in 59 of 60 runs; competitor A in 57; competitor B in 28. Repetition is what turns a lucky answer into a measurement.

## Why measure share of voice instead of an 'AI rank'?

Because answers are non-deterministic — near-zero odds of identical brand lists across repeated runs — so a single-run 'rank' measures the run, not the brand. Rates across samples are stable enough to act on and honest enough to publish. When an answer happens to be a ranked list, the position within it is worth recording as an observed fact; it just isn't the headline number.

[How SeeGeo tracks it](https://see-geo.com/about)

## How do you actually grow share of voice?

Three compounding levers, none fast. On your site: extractability, so the passages engines lift can be yours. Off it: presence on the comparison pages and directories answers are assembled from. Underneath both: entity consistency, so systems are confident enough about what you are to name you. Expect movement over weeks, not days — which is why the measurement has to be a stable panel run on a schedule.

## Frequently asked questions

### What does Share of voice (AI) mean?

In AI visibility, share of voice is how often you are named in AI answers for your category's questions relative to your competitors — the answer-engine equivalent of market share of attention.

### What is a good AI share of voice?

Benchmarks are young, but the working thresholds: appearing in a majority of category answers puts you in the conversation; parity with your strongest competitor means the answers treat you as a peer; dominance is being the default first name. Direction over weeks matters more than any absolute number.

### How many prompts and runs make the measurement honest?

Enough repetition to separate signal from answer-to-answer noise — in practice panels of 20+ questions with 3+ runs each. One run of one prompt is an anecdote; the trend of a repeated panel is a metric.

### Can share of voice be high while citations are zero?

Yes, and it's common: mentions flow from training data while citations require your pages to be retrieved live. A brand can be named in nearly every answer whose sources never include its own site — visibility built entirely on third-party pages.

## Related pages

- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility)
- [What is Prompt panel?](https://see-geo.com/glossary/prompt-panel)
- [What is AI citation?](https://see-geo.com/glossary/ai-citation)

---

# What is Prompt panel?
Source: https://see-geo.com/glossary/prompt-panel · Updated 2026-09-08

A prompt panel is the fixed set of realistic customer questions used to measure AI visibility — the questionnaire you re-ask the engines on a schedule so results are comparable over time.

## What does a good prompt panel contain?

Four intents, weighted toward the first: unbranded category questions ('best issue tracker for a small team') where visibility is won or lost; head-to-head comparisons ('X vs Y'); use-case questions phrased the way buyers actually ask; and a few branded questions as a floor check — if an engine can't answer 'what is X?', nothing else matters.

## Why must the panel stay fixed?

Because the metric is the trend. Change the questions and the numbers move for reasons that have nothing to do with your visibility, which silently corrupts every comparison to last month. Add questions deliberately and version the panel; never swap quietly.

[Share of voice, explained](https://see-geo.com/glossary/share-of-voice)

## What's the most common panel mistake?

The vanity panel: mostly branded questions, which score beautifully — engines can usually answer 'what is [your name]?' — while measuring nothing contested. A panel earns its keep in the unbranded majority, where a stranger's question either surfaces you or surfaces your competitor. If your panel's numbers look flattering, that's usually a description of the panel, not of your visibility.

## Frequently asked questions

### What does Prompt panel mean?

A prompt panel is the fixed set of realistic customer questions used to measure AI visibility — the questionnaire you re-ask the engines on a schedule so results are comparable over time.

### How many prompts should a panel have?

Twenty to forty covers most small businesses: enough spread across intents to be representative, small enough to re-run frequently without the cost mattering. Bigger panels buy precision mostly for brands competing across many product lines.

### Should prompts include my brand name?

A few should — they establish the floor of whether engines know you exist. But most of the panel belongs to unbranded category questions, because that's where a customer who hasn't heard of you either finds you or finds your competitor.

### How often should a panel run?

Weekly is the useful default: frequent enough to catch shifts, spaced enough that each run is cheap and the trend line means something. Daily runs add value mainly when you're actively shipping changes and want fast feedback.

## Related pages

- [What is Share of voice (AI)?](https://see-geo.com/glossary/share-of-voice)
- [What is AI visibility?](https://see-geo.com/glossary/ai-visibility)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)

---

# What is Content extractability?
Source: https://see-geo.com/glossary/content-extractability · Updated 2026-09-08

Content extractability is how easily an AI system can lift a clear, self-contained statement from a page — the property that decides whether the page gets quoted in answers or merely read and discarded.

## What makes content extractable?

Passages that stand alone. An extractable page answers its core question in the first two paragraphs, uses headings shaped like the questions people ask, states facts with numbers and named sources, and carries visible dates so machines can place it in time. Each of those is a shape a model can quote without reconstructing meaning from context.

- Answer-first structure: the direct answer before the background, not after.
- Question-shaped headings that mirror real queries.
- Statistics with named sources — the highest-leverage single edit measured (30–40% visibility gains, Princeton GEO study, KDD 2024).
- Visible 'updated' dates, honest ones.
- Comparisons in tables and lists rather than buried in prose.

## How is extractability measured?

By scoring the shapes directly: does the opening answer the page's question, are headings question-shaped, do statistics and sources appear, is there a visible date, is comparative content structured? SeeGeo's audit scores each and shows the per-page breakdown — the same rubric this glossary is written to pass.

[Score your pages free](https://see-geo.com/)

## What's the most common extractability mistake?

The throat-clearing introduction. Pages that open with three paragraphs of scene-setting — the kind that begins 'In today's fast-paced digital landscape' — push the actual answer below where retrieval looks for it. The test is brutal and useful: delete your first three paragraphs and see if the page got better. On most of the web, it does.

## Frequently asked questions

### What does Content extractability mean?

Content extractability is how easily an AI system can lift a clear, self-contained statement from a page — the property that decides whether the page gets quoted in answers or merely read and discarded.

### Is extractability just 'writing well'?

No — plenty of excellent prose is unextractable because its meaning accumulates across paragraphs. Extractability is a structural property: whether single passages survive being lifted out alone. A mediocre page with a crisp definition often outperforms an elegant essay.

### Does extractability help classic SEO too?

Substantially — the same shapes win featured snippets and AI Overview citations, and readers scan the same way machines lift. It's the rare optimization with no trade-off against human readers.

### What's the fastest extractability win?

Add a two-sentence direct answer at the top of each important page, under a heading phrased as the question. It's an afternoon of editing and it targets the exact passage an engine looks for first.

## Related pages

- [What is Answer-first content?](https://see-geo.com/glossary/answer-first-content)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data)

---

# What is Entity clarity?
Source: https://see-geo.com/glossary/entity-clarity · Updated 2026-09-08

Entity clarity is how unambiguously machines can determine what a business is — its name, category, location, and offering — from its website and the public records around it.

## What signals establish an entity?

On-site: a plain-prose sentence saying what you are (not only in JSON-LD — machines that read rendered text, and humans, need it too), Organization structured data with a reachable contact, an About page, and consistent naming everywhere. Off-site: presence in public registries like Wikidata and consistency across the directories and profiles that mention you. Ambiguity anywhere makes systems hedge — and a hedging system recommends someone clearer.

[What is structured data?](https://see-geo.com/glossary/structured-data)

## Why do AI systems weigh entity clarity so heavily?

Because recommending a business is an act of trust: the system must be confident the thing exists, is what it claims, and is distinguishable from similarly named things. SeeGeo's own audit initially scored its maker 66/100 here — the definition lived only in structured data, with no About page — and an afternoon of exactly these fixes moved it to 92. The levers are cheap; the neglect is just common.

[That case study](https://see-geo.com/blog/we-audited-ourselves)

## What's the most common entity mistake?

Inconsistency across surfaces: one name on the site, a variant on the directories, a third in old profiles, each with a slightly different description. Machines resolve entities by matching signals across sources — every mismatch lowers their confidence that these are the same thing, and a low-confidence entity gets hedged out of recommendations. One canonical sentence, used identically everywhere, beats ten artisanal variants.

## Frequently asked questions

### What does Entity clarity mean?

Entity clarity is how unambiguously machines can determine what a business is — its name, category, location, and offering — from its website and the public records around it.

### What's the fastest entity-clarity fix?

Write one sentence — '[Name] is a [category] that [what it does] for [whom]' — and put it in your homepage's visible copy, your About page, and your Organization schema, identically. Consistency across the three is itself the signal.

### Does an About page really matter to machines?

Yes — it's a canonical place both search quality guidelines and AI source-weighing look for who's behind a site. Its absence reads as a trust gap; a factual one with real contact details anchors the entity cheaply.

### What is Wikidata and should I be in it?

A public database of entities that many AI systems consult to confirm a thing exists and what it is. A minimal factual item — name, site, type, location — takes about fifteen minutes and is a legitimate, free credibility anchor.

## Related pages

- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data)
- [What is E-E-A-T?](https://see-geo.com/glossary/e-e-a-t)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)

---

# What is Structured data (Schema.org)?
Source: https://see-geo.com/glossary/structured-data · Updated 2026-09-08

Structured data is machine-readable markup — usually JSON-LD using Schema.org vocabulary — that tells crawlers explicitly what a page contains: an organization, a product, an FAQ, an article.

## Which schema types matter most for AI visibility?

Organization (who you are, with a reachable contact), WebSite/WebApplication (what the site is), FAQPage (packages Q&A in exactly the liftable shape answers use), Article with real authorship, and where honest, Product and LocalBusiness. Linked together with stable @id references they form one graph a machine can walk instead of islands it must guess about.

## What's the one rule of structured data?

Never mark up what the visible page doesn't say. Schema describing invisible content — or worse, invented ratings — is the pattern search engines penalize and the exact dishonesty AI systems are learning to discount. The test is symmetry: a human reading the page and a machine reading the markup should come away with the same facts.

[Entity clarity, explained](https://see-geo.com/glossary/entity-clarity)

## How much structured data is enough?

Less than the plugins suggest. The core set — Organization, WebSite, the page's own type, FAQPage where genuine Q&A exists — covers what engines actually consume; stacking every applicable type adds maintenance surface without visibility. Volume is not the axis that matters: one honest, valid, interlinked graph outperforms twelve types of markup that drift from the visible page.

## Frequently asked questions

### What does Structured data (Schema.org) mean?

Structured data is machine-readable markup — usually JSON-LD using Schema.org vocabulary — that tells crawlers explicitly what a page contains: an organization, a product, an FAQ, an article.

### Does structured data directly improve AI visibility?

Indirectly but meaningfully: it removes ambiguity about what you are, which feeds entity confidence, and FAQPage markup hands engines pre-packaged liftable answers. It complements visible content — it never substitutes for it.

### JSON-LD, microdata, or RDFa?

JSON-LD — it's Google's stated preference, lives in one script tag instead of woven through your HTML, and is the format AI toolchains parse most reliably. There's no modern reason to choose the others.

### How do I check my structured data is valid?

Google's Rich Results Test and Schema.org's validator catch syntax; the harder check is honesty — whether the markup matches the visible page. SeeGeo's audit tests both, including whether your schema description has a visible-prose counterpart.

## Related pages

- [What is Entity clarity?](https://see-geo.com/glossary/entity-clarity)
- [What is Content extractability?](https://see-geo.com/glossary/content-extractability)
- [What is E-E-A-T?](https://see-geo.com/glossary/e-e-a-t)

---

# What is robots.txt?
Source: https://see-geo.com/glossary/robots-txt · Updated 2026-09-08

robots.txt is a plain-text file at a website's root that tells crawlers which parts of the site they may read, using the Robots Exclusion Protocol (RFC 9309) — and it's where AI crawler access is granted or denied.

## How does robots.txt control AI visibility?

Every major AI crawler — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — checks robots.txt before reading and honors what it finds. One leftover 'Disallow: /' under a wildcard user-agent, written years ago for some other reason, silently removes a site from AI training data, AI search indexes, and live retrieval at once. The file is tiny; its blast radius is not.

```
# Allow a specific AI crawler
User-agent: GPTBot
Allow: /

# Block one, allow the rest
User-agent: CCBot
Disallow: /
```

[Every crawler's exact lines](https://see-geo.com/bots)

## What can't robots.txt control?

The network layer. CDNs and firewalls decide whether a crawler's request ever reaches your server — Cloudflare blocks AI crawlers by default for many accounts — and no robots.txt directive can override a blocked connection. Auditing access means checking both layers, which is why SeeGeo tests each crawler's real reachability rather than reading your rules and assuming.

## What's the most common robots.txt mistake?

The leftover blanket block: a 'User-agent: * / Disallow: /' written for a staging site and shipped to production, or an old rule blocking a directory that modern crawlers need. Because the file fails silently — nothing breaks, you just quietly vanish — these survive for years. Re-check it after every migration, and test what crawlers actually experience rather than what the file appears to say.

## Frequently asked questions

### What does robots.txt mean?

robots.txt is a plain-text file at a website's root that tells crawlers which parts of the site they may read, using the Robots Exclusion Protocol (RFC 9309) — and it's where AI crawler access is granted or denied.

### Do AI companies actually respect robots.txt?

The major operators — OpenAI, Anthropic, Google, Perplexity — publicly commit to it and observably comply. Edge cases exist among smaller scrapers, but the dominant real-world problem is the reverse: sites blocking legitimate crawlers by accident.

### What does 'User-agent: *' do to AI crawlers?

The wildcard applies to every crawler without a more specific rule — including all AI crawlers. A 'Disallow: /' under it blocks everything from everyone, which is the single most common accidental way sites vanish from AI answers.

### Where do I put robots.txt?

At the exact root: yoursite.com/robots.txt. Crawlers look only there — a robots.txt in a subdirectory is ignored entirely, and a missing file means no rules, which permits everything.

## Related pages

- [What is AI crawler?](https://see-geo.com/glossary/ai-crawler)
- [What is LLM training data?](https://see-geo.com/glossary/llm-training-data)
- [What is Google AI Overviews?](https://see-geo.com/glossary/ai-overviews)

---

# What is LLM training data?
Source: https://see-geo.com/glossary/llm-training-data · Updated 2026-09-08

LLM training data is the text corpus a language model learns from — largely crawled web content — and it determines what the model 'knows' about your business when answering without live search.

## How does your website end up in training data?

Through training crawlers — GPTBot, ClaudeBot, Google-Extended — and public archives like Common Crawl (collected by CCBot), which many labs train on. What they collected months or years ago shapes today's ungrounded answers; what they collect today shapes the next model generation. Training-data visibility is an investment with a long settlement date.

[CCBot and Common Crawl, explained](https://see-geo.com/bots/ccbot)

## Should you block training crawlers?

It's the one genuinely two-sided crawler decision. Blocking keeps your content out of future models — a real right, rationally exercised by publishers whose content is their product. Allowing means future models learn your business exists, matters when customers ask them for recommendations, and costs nothing today. For businesses that want to be found, the visibility case usually wins; for content businesses, protection often does.

## Can you get removed from training data retroactively?

Effectively no. Blocking training crawlers is forward-only — models already trained keep what they learned, and there is no general mechanism to unlearn a website from a shipped model. Vendor-specific removal processes exist but are slow and partial. Which cuts both ways: the visibility you build in training data is similarly durable, compounding quietly across model generations.

## Frequently asked questions

### What does LLM training data mean?

LLM training data is the text corpus a language model learns from — largely crawled web content — and it determines what the model 'knows' about your business when answering without live search.

### If I block GPTBot today, does ChatGPT forget me?

No — models already trained retain what they learned, and blocking doesn't reach back. It stops your content entering future training runs, with effects that appear only when those future models ship.

### Why does ChatGPT know outdated facts about my business?

Its memory is a snapshot from when its training data was collected. Ungrounded answers speak from that snapshot; only grounded answers (with live search) can see your current site. Fixing the record means both updating your site and being visible enough that grounded retrieval finds the update.

### Is Common Crawl the same as Google's index?

No — Common Crawl is a nonprofit public web archive many AI labs train on, collected by CCBot. Google's index is proprietary and gathered by Googlebot. Blocking one has no effect on the other.

## Related pages

- [What is AI crawler?](https://see-geo.com/glossary/ai-crawler)
- [What is Grounding?](https://see-geo.com/glossary/grounding)
- [What is robots.txt?](https://see-geo.com/glossary/robots-txt)

---

# What is E-E-A-T?
Source: https://see-geo.com/glossary/e-e-a-t · Updated 2026-09-08

E-E-A-T — Experience, Expertise, Authoritativeness, Trustworthiness — is Google's framework for judging content quality, and increasingly the checklist AI systems apply when deciding which sources to trust and cite.

## What does E-E-A-T look like to a machine?

Checkable signals: named authors with stated credentials rather than anonymous bylines, an About page saying who is behind the site, reachable contact details, citations to named sources, honest dates, and claims that survive verification. Each is mechanically detectable, which is exactly why both search raters and AI source-selection lean on them: they're cheap proxies for 'someone accountable stands behind this'.

## Does E-E-A-T apply to AI answers too?

Increasingly, yes. Answer engines choosing between candidate sources favor the same trust shape — named authorship, verifiable claims, institutional presence. Anonymous content farms are precisely what their source-quality filters exist to exclude, and 'by Team' bylines sit closer to that pattern than most companies would like.

[Entity clarity, explained](https://see-geo.com/glossary/entity-clarity)

## What's the most common E-E-A-T mistake?

Faking it: stock-photo 'team members', invented credentials, testimonial markup for reviews that don't exist. Fabricated trust signals are worse than absent ones — they're the exact pattern quality systems are trained to catch, and one detected fake discounts everything else you claim. Verifiable minimalism (a real name, a real inbox, honest dates) beats fabricated depth every time.

## Frequently asked questions

### What does E-E-A-T mean?

E-E-A-T — Experience, Expertise, Authoritativeness, Trustworthiness — is Google's framework for judging content quality, and increasingly the checklist AI systems apply when deciding which sources to trust and cite.

### Is E-E-A-T a Google ranking factor?

Not a single scored factor — it's the rubric human quality raters apply, which trains the systems that do rank. The practical effect is the same: content exhibiting the signals outperforms content that hides its authorship.

### What's the cheapest E-E-A-T improvement?

Real bylines: a person's name and a one-line credential on each substantial piece of content. It costs nothing, is verifiable, and moves the exact signal both raters and AI source-selection check first.

### Does E-E-A-T matter for small businesses?

More than for big ones, relatively: a small business can't outspend anyone on content volume, but named humans, real experience, and checkable claims are trust signals available at any size — and they're what distinguish you from the anonymous content competing for the same answers.

## Related pages

- [What is Entity clarity?](https://see-geo.com/glossary/entity-clarity)
- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data)
- [What is AI citation?](https://see-geo.com/glossary/ai-citation)

---

# What is Answer-first content?
Source: https://see-geo.com/glossary/answer-first-content · Updated 2026-09-08

Answer-first content states its complete answer in the opening one or two paragraphs — before background, before caveats — so the direct answer is the first thing both readers and AI systems encounter.

## Why does answer-first structure win AI visibility?

Because retrieval looks for a passage that resolves the question on its own, and the opening is where it looks first. Pages that build context before concluding force the engine to reconstruct the answer across paragraphs — which it often won't do when a competitor's page hands it a clean opening statement. The inverted pyramid, a century old in journalism, turns out to be the native shape of AI-quotable content.

## How do you retrofit answer-first structure?

For each important page, write the two-sentence complete answer to the page's core question and move it to the top, bolded, under a question-shaped heading. Keep the depth below — answer-first doesn't mean answer-only; it means the depth supports a stated conclusion instead of delaying it.

[Content extractability, explained](https://see-geo.com/glossary/content-extractability) · [The GEO writing checklist](https://see-geo.com/blog/geo-writing-checklist-get-cited-by-chatgpt)

## What does answer-first look like in practice?

Open any important page and read only its first two paragraphs. If a stranger would come away with the actual answer — not the promise of one — the page is answer-first. The classic failure is the SEO-era intro that circles the topic to accumulate keywords before committing to a claim; engines skip it, and increasingly, so do readers arriving from an AI answer that already summarized the circling.

## Frequently asked questions

### What does Answer-first content mean?

Answer-first content states its complete answer in the opening one or two paragraphs — before background, before caveats — so the direct answer is the first thing both readers and AI systems encounter.

### Doesn't giving the answer up front kill reader engagement?

The evidence says the opposite: readers who get the answer immediately trust the page and stay for the depth, while readers made to scroll for it bounce to a page that respects their time. Answer-first is a reader courtesy that happens to also be machine-optimal.

### Is answer-first the same as a TL;DR?

Close, with one difference: a TL;DR summarizes the page, while an answer-first lead answers the question — self-contained, no 'as we'll see below'. The lead should survive being quoted entirely on its own, because that's precisely how an engine will use it.

### Where else does the answer belong besides the top?

In the FAQ section as a compact restatement, and in FAQPage structured data — the three placements reinforce each other, giving engines the same answer in prose, in Q&A shape, and in markup.

## Related pages

- [What is Content extractability?](https://see-geo.com/glossary/content-extractability)
- [What is Generative Engine Optimization (GEO)?](https://see-geo.com/glossary/generative-engine-optimization)
- [What is Structured data (Schema.org)?](https://see-geo.com/glossary/structured-data)

---

# Unblock AI crawlers: Cloudflare, Akamai, DataDome, Imperva
Source: https://see-geo.com/unblock · Updated 2026-09-02

Your firewall may be hiding you from ChatGPT, Claude and Perplexity. Vendor guides to allow AI search crawlers without opening the door to scrapers.

## Pages

- [How do I unblock AI crawlers on Cloudflare?](https://see-geo.com/unblock/cloudflare) — Cloudflare's 'Block AI bots' toggle and Super Bot Fight Mode hide sites from ChatGPT, Claude and Perplexity. The settings to change, and how to verify.
- [How do I unblock AI crawlers on Akamai?](https://see-geo.com/unblock/akamai) — Akamai Bot Manager denies or challenges AI crawlers by category. Which categories to set to Allow, how to define missing bots, and how to verify.
- [How do I unblock AI crawlers on DataDome?](https://see-geo.com/unblock/datadome) — DataDome challenges any client it can't identify as a known good bot, AI crawlers included. Which categories to allow, the custom rule, and how to verify.
- [How do I unblock AI crawlers on Imperva?](https://see-geo.com/unblock/imperva) — Imperva (Incapsula) challenges suspected bots and blocks unknown ones, hiding you from AI assistants. The settings to change, and how to verify.

---

# How do I unblock AI crawlers on Cloudflare?
Source: https://see-geo.com/unblock/cloudflare · Updated 2026-09-02

On Cloudflare, AI crawlers are usually blocked by one of three things: the one-click "Block AI bots" toggle, Super Bot Fight Mode challenging "definitely automated" traffic, or a WAF custom rule — and the fix is to allow verified search and AI crawlers in AI Crawl Control while leaving training crawlers to your own policy. Ten minutes in the dashboard, then verify with a crawler-user-agent request.

## Why is Cloudflare blocking AI crawlers on my site?

Cloudflare made blocking AI crawlers a single switch, available on every plan including free, and many owners flipped it to stop model training without realizing it also blocks the crawlers that answer customer questions. Separately, Bot Fight Mode and Super Bot Fight Mode challenge any client Cloudflare classifies as automated — AI crawlers included — with a JavaScript challenge that no crawler can solve. The challenge page is what an AI assistant receives instead of your homepage, and it is also what SeeGeo's audit received when it graded the site as walled.

Since 2025 Cloudflare has also shipped a managed robots.txt with Content Signals (search / ai-input / ai-train directives) and an AI Crawl Control panel with per-crawler allow and block decisions. Those are the right tools for a deliberate policy; the problem is defaults that were never a decision.

## Where does the block live in the Cloudflare dashboard?

Check these in order; the first one you find is usually the whole story. Menu names are current as of mid-2026 and shift occasionally.

- Security → Bots → "Block AI bots" — the one-click toggle. On means every known AI crawler, search ones included, is blocked at the edge.
- Security → Bots → Super Bot Fight Mode (or Bot Fight Mode on free plans) → "Definitely automated" set to Block or Managed Challenge, and "Allow verified bots" turned off.
- Security → AI Crawl Control — the per-crawler table (OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot, Google-Extended…) with Allow / Block per row.
- Security → WAF → Custom rules — a rule matching user-agents ("bot", "crawler", "GPT") or ASNs, or a rate limit tight enough to catch a polite crawl.
- Your robots.txt — Cloudflare's managed robots.txt can add Content Signals and a block on AI training crawlers; check what it actually says at yoursite.com/robots.txt.

## How do I allow AI crawlers on Cloudflare, step by step?

Do these top to bottom; each is reversible and takes effect within a minute.

- 1. Security → Bots: turn "Block AI bots" OFF. If you want to keep blocking training crawlers, do that per-crawler in the next step instead of with this blanket switch.
- 2. Security → AI Crawl Control: set OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot and Googlebot to Allow. Decide GPTBot and Google-Extended (training) on their own merits — allowing them is not required to be cited.
- 3. Security → Bots → Super Bot Fight Mode: turn "Allow verified bots" ON. OpenAI's, Anthropic's, Perplexity's and Google's crawlers are on Cloudflare's verified list, so this admits them while still challenging unverified automation.
- 4. If "Definitely automated" is set to Block, either change it to Allow or add a WAF custom rule that Skips bot protection for verified bots — expression: (cf.verified_bot_category in {"Search Engine Crawler" "AI Crawler"}) with action Skip → Super Bot Fight Mode.
- 5. Security → WAF → Custom rules: read every rule that mentions user-agent, "bot" or "crawler" and add an exception for the crawlers above, or delete the rule if it was a default you never chose.
- 6. Open yoursite.com/robots.txt. If it contains Content-Signal lines or Disallow rules for search crawlers you want, edit them under Security → Bots → Managed robots.txt (or your origin's robots.txt if managed mode is off). Keep ai-train=no if that's your policy; make sure search=yes.
- 7. Verify with the crawler-user-agent requests below, then re-run the audit.

## Which AI crawlers should I allow, and which can I keep blocking?

The distinction that matters is search versus training. Search crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Googlebot) are what makes an assistant able to find and cite you; blocking them makes you invisible in AI answers. Training crawlers (GPTBot, Google-Extended, CCBot) feed model training and blocking them costs you nothing in visibility today. Most "Cloudflare blocks AI bots" settings treat both groups as one, which is exactly why owners who only meant to opt out of training end up invisible.

| Crawler | What it feeds | Allow? |
|---|---|---|
| OAI-SearchBot | ChatGPT search answers and citations | Yes — this is the one that recommends you |
| ChatGPT-User | Live fetches when a user asks ChatGPT about a page | Yes |
| GPTBot | OpenAI model training | Your call — no effect on being cited today |
| ClaudeBot / Claude-User | Anthropic's index and live fetches for Claude | Yes |
| PerplexityBot / Perplexity-User | Perplexity answers and citations | Yes |
| Googlebot | Google Search AND AI Overviews / AI Mode | Yes — blocking it removes you from Google entirely |
| Google-Extended | Gemini training (not Search) | Your call |
| Bytespider, CCBot | Third-party scrapers and training sets | Block if you like — no visibility cost |

[Every crawler, one page each: what it is and how to control it](https://see-geo.com/bots) · [The platform toggles that block the wrong crawlers](https://see-geo.com/blog/website-platform-ai-visibility-defaults)

## How do I verify the wall is actually open?

Test from outside, as a crawler would — not from your browser, which is exactly the client the wall was built to admit. Run these from any terminal (or an online HTTP tester) and compare the responses:

A healthy answer is a 200 status with your real HTML. A 403, a challenge page, or a response with the vendor's mitigation header means the crawler is still blocked. If your wall verifies bots by IP range rather than user-agent, a spoofed user-agent from your laptop may still be challenged even though the real crawler gets through — in that case the definitive test is the vendor's own bot analytics, or simply re-running the audit and checking the crawler table.

```
# As ChatGPT's search crawler:
curl -sI -A "OAI-SearchBot/1.0" https://yoursite.com/ | head -5
# As Claude's crawler:
curl -sI -A "ClaudeBot/1.0" https://yoursite.com/ | head -5
# As a plain browser, for comparison:
curl -sI -A "Mozilla/5.0" https://yoursite.com/ | head -5
# SeeGeo's own crawler, if you want the audit itself to get through:
curl -sI -A "SeeGeoAudit/1.0" https://yoursite.com/ | head -5
```

[Re-run the free audit — the crawler table is the receipt](https://see-geo.com/)

## Frequently asked questions

### Does Cloudflare's 'Block AI bots' toggle block ChatGPT search too?

Yes. The one-click toggle blocks the known AI crawlers as a group, and that group includes OAI-SearchBot and ChatGPT-User — the crawlers that make ChatGPT able to cite you — not just GPTBot, the training crawler. To block training while staying citable, turn the toggle off and use AI Crawl Control to decide per crawler.

### What does the cf-mitigated: challenge header mean?

It's Cloudflare telling the client it was served a challenge page instead of your site. When a crawler's request comes back with that header, the crawler saw a JavaScript puzzle, not your content — which for an AI assistant is the same as the page being blank. It's also the signal SeeGeo's audit uses to report a Cloudflare wall.

### Will allowing verified bots let scrapers in?

No. Cloudflare's verified-bot list is a vetted registry of crawlers that prove their identity through published IP ranges or reverse DNS; scrapers spoofing a Googlebot user-agent don't qualify and stay challenged. Allowing verified bots is the narrowest change that admits real search and AI crawlers.

### Do I need to allow GPTBot to show up in ChatGPT?

No. GPTBot collects training data; OAI-SearchBot and ChatGPT-User are what ChatGPT's search and browsing use to find and cite pages. You can block GPTBot and still be cited, as long as the search crawlers are allowed.

## Related pages

- [How do I unblock AI crawlers on Akamai?](https://see-geo.com/unblock/akamai)
- [How do I unblock AI crawlers on DataDome?](https://see-geo.com/unblock/datadome)
- [How do I unblock AI crawlers on Imperva?](https://see-geo.com/unblock/imperva)

---

# How do I unblock AI crawlers on Akamai?
Source: https://see-geo.com/unblock/akamai · Updated 2026-09-02

On Akamai, AI crawlers are blocked when Bot Manager's category action for AI or LLM bots is Deny, when an unknown bot's default action is Deny or Challenge, or when a WAF rule matches the user-agent — and the fix is to set the search-engine and AI-bot categories to Allow or Monitor, define any crawler Akamai doesn't yet categorize, and activate the configuration. Expect a few minutes for activation to reach production.

## Why is Akamai blocking AI crawlers on my site?

Akamai Bot Manager sorts traffic into Akamai-categorized bots (search engines, monitoring tools, and since 2024–2025 AI and large-language-model crawlers), customer-defined bots, and everything else it detects as automated. Each category carries an action — Allow, Monitor, Deny, or a challenge — and enterprise defaults lean toward Deny for anything not explicitly trusted. A crawler that arrives outside a trusted category gets a 403 with an Akamai reference number instead of your page; that reference page is what SeeGeo's audit received when it reported an Akamai wall.

Akamai also offers Content Protector and App & API Protector rules that can match crawler user-agents or rate-limit polite crawls. Because Akamai changes ship through a staged activation, a block that was added months ago can outlive whoever added it.

## Where does the block live in Akamai Control Center?

Names below are current as of mid-2026; your account's security configuration may label the screens slightly differently.

- Security Configuration → Bot Manager → Akamai-Categorized Bots: the per-category action table. Look for "Search Engine Bots" and the AI / LLM crawler category (Akamai has named it "AI Bots" or "Large Language Model Crawlers" across releases).
- Bot Manager → Bot Detection (or "Unknown bots"): the default action for automated traffic Akamai can't categorize — a Deny here catches newer crawlers before they're categorized.
- Bot Manager → Custom Bot Categories / Customer-Defined Bots: where you define a crawler by user-agent string when Akamai hasn't categorized it yet.
- App & API Protector / Kona Site Defender → Custom rules: user-agent matches, rate controls and Client Reputation thresholds that can catch crawlers indiscriminately.

## How do I allow AI crawlers on Akamai, step by step?

Every change below lands only after you activate the configuration — save, activate to staging, test, then activate to production.

- 1. Bot Manager → Akamai-Categorized Bots: set "Search Engine Bots" to Allow (Googlebot, Bingbot) and the AI / LLM crawler category to Monitor or Allow. Monitor lets you watch the traffic for a week before committing to Allow.
- 2. For crawlers missing from Akamai's categories — check for OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot — create a customer-defined bot per user-agent under Custom Bot Categories, and give the category the Allow or Monitor action.
- 3. Bot Detection: confirm the action for unknown automated traffic isn't silently denying the crawlers above; if it is, either raise the AI category above it or set the detection to Monitor for those user-agents.
- 4. App & API Protector custom rules: search rule conditions for "bot", "crawler", "GPT", "Claude" and remove or exempt the AI crawlers; check rate controls aren't tighter than a polite one-request-per-second crawl.
- 5. Save → activate to staging → verify from staging with the crawler-user-agent requests below → activate to production.
- 6. Re-run the audit after production activation and confirm the crawler table shows the AI crawlers allowed.

## Which AI crawlers should I allow, and which can I keep blocking?

The distinction that matters is search versus training. Search crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Googlebot) are what makes an assistant able to find and cite you; blocking them makes you invisible in AI answers. Training crawlers (GPTBot, Google-Extended, CCBot) feed model training and blocking them costs you nothing in visibility today. Most "Akamai blocks AI bots" settings treat both groups as one, which is exactly why owners who only meant to opt out of training end up invisible.

| Crawler | What it feeds | Allow? |
|---|---|---|
| OAI-SearchBot | ChatGPT search answers and citations | Yes — this is the one that recommends you |
| ChatGPT-User | Live fetches when a user asks ChatGPT about a page | Yes |
| GPTBot | OpenAI model training | Your call — no effect on being cited today |
| ClaudeBot / Claude-User | Anthropic's index and live fetches for Claude | Yes |
| PerplexityBot / Perplexity-User | Perplexity answers and citations | Yes |
| Googlebot | Google Search AND AI Overviews / AI Mode | Yes — blocking it removes you from Google entirely |
| Google-Extended | Gemini training (not Search) | Your call |
| Bytespider, CCBot | Third-party scrapers and training sets | Block if you like — no visibility cost |

[Every crawler, one page each: what it is and how to control it](https://see-geo.com/bots) · [The platform toggles that block the wrong crawlers](https://see-geo.com/blog/website-platform-ai-visibility-defaults)

## How do I verify the wall is actually open?

Test from outside, as a crawler would — not from your browser, which is exactly the client the wall was built to admit. Run these from any terminal (or an online HTTP tester) and compare the responses:

A healthy answer is a 200 status with your real HTML. A 403, a challenge page, or a response with the vendor's mitigation header means the crawler is still blocked. If your wall verifies bots by IP range rather than user-agent, a spoofed user-agent from your laptop may still be challenged even though the real crawler gets through — in that case the definitive test is the vendor's own bot analytics, or simply re-running the audit and checking the crawler table.

```
# As ChatGPT's search crawler:
curl -sI -A "OAI-SearchBot/1.0" https://yoursite.com/ | head -5
# As Claude's crawler:
curl -sI -A "ClaudeBot/1.0" https://yoursite.com/ | head -5
# As a plain browser, for comparison:
curl -sI -A "Mozilla/5.0" https://yoursite.com/ | head -5
# SeeGeo's own crawler, if you want the audit itself to get through:
curl -sI -A "SeeGeoAudit/1.0" https://yoursite.com/ | head -5
```

[Re-run the free audit — the crawler table is the receipt](https://see-geo.com/)

## Frequently asked questions

### Does Akamai have an AI crawler category, or do I define each bot myself?

Both, depending on your account's release: Akamai added AI and LLM crawler categories to Bot Manager's categorized bots in 2024–2025, so most accounts can set one category action. Newer crawlers can lag categorization, in which case a customer-defined bot by user-agent covers them until Akamai catches up.

### Why does my Akamai site return a reference number page to crawlers?

That page ("Access Denied — You don't have permission… Reference #18.xxx") is Akamai's Deny action. The reference number is the incident ID your Akamai support team can look up to tell you which rule fired — the fastest way to find the block if the category table looks fine.

### How long until an Akamai change takes effect?

Akamai security configurations must be activated; staging activation is typically a few minutes and production a few more. Changes saved but not activated do nothing, which is the most common reason a fix 'didn't work' — check the activation status before assuming the setting was wrong.

## Related pages

- [How do I unblock AI crawlers on Cloudflare?](https://see-geo.com/unblock/cloudflare)
- [How do I unblock AI crawlers on DataDome?](https://see-geo.com/unblock/datadome)
- [How do I unblock AI crawlers on Imperva?](https://see-geo.com/unblock/imperva)

---

# How do I unblock AI crawlers on DataDome?
Source: https://see-geo.com/unblock/datadome · Updated 2026-09-02

DataDome blocks AI crawlers when the AI-crawler category in its known-bots settings is set to block, or when a crawler isn't recognized and hits the default protection that challenges every unidentified client — and the fix is to allow the search-engine and AI-assistant categories, add a custom allow rule for any crawler DataDome doesn't recognize, and confirm the endpoints you care about aren't in a stricter mode.

## Why is DataDome blocking AI crawlers on my site?

DataDome's model is strict by design: traffic is either a recognized good bot, a human that passes its JavaScript and behavioral checks, or it's challenged. A crawler can't run the challenge, so an unrecognized crawler receives a 403 carrying the x-datadome header and a captcha interstitial instead of your page — that header is what SeeGeo's audit detects when it reports a DataDome wall.

DataDome also introduced explicit AI-crawler handling in 2024–2025, separating AI assistants and search crawlers from AI data scrapers, with account-level defaults that often block the whole group. Owners who wanted to stop scraping frequently ended up blocking the assistants too.

## Where does the block live in the DataDome dashboard?

Labels are current as of mid-2026 and DataDome renames screens periodically; look for the concepts rather than the exact words.

- Dashboard → Management → Bot settings / Known bots (sometimes "Verified bots"): categories such as Search Engines, Monitoring, AI Assistants / AI Crawlers, and AI Data Scrapers, each with allow or block.
- Dashboard → Management → Custom rules: allow, block or challenge by user-agent, IP, ASN, endpoint or country — the place to add a crawler DataDome doesn't categorize.
- Dashboard → Management → Endpoints (Protection vs Monitor): a page in Protection mode challenges everything unrecognized; Monitor mode only records.
- Your robots.txt, which DataDome doesn't manage — check it separately.

## How do I allow AI crawlers on DataDome, step by step?

Changes apply in real time; there is no activation step.

- 1. Known bots: confirm Search Engines is allowed (it usually is) and set the AI Assistants / AI Crawlers category to Allow. Leave AI Data Scrapers on block if scraping was the problem you were solving.
- 2. For any crawler that isn't in a category — check OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot — add a custom rule: condition user-agent contains the crawler name, action Allow, scope your whole site.
- 3. Endpoints: confirm your homepage and key content pages aren't under a stricter per-endpoint rule than the site default.
- 4. Verify with the crawler-user-agent requests below; a response without the x-datadome header and without a captcha page means the crawler is through.
- 5. Re-run the audit and check the crawler table.

## Which AI crawlers should I allow, and which can I keep blocking?

The distinction that matters is search versus training. Search crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Googlebot) are what makes an assistant able to find and cite you; blocking them makes you invisible in AI answers. Training crawlers (GPTBot, Google-Extended, CCBot) feed model training and blocking them costs you nothing in visibility today. Most "DataDome blocks AI bots" settings treat both groups as one, which is exactly why owners who only meant to opt out of training end up invisible.

| Crawler | What it feeds | Allow? |
|---|---|---|
| OAI-SearchBot | ChatGPT search answers and citations | Yes — this is the one that recommends you |
| ChatGPT-User | Live fetches when a user asks ChatGPT about a page | Yes |
| GPTBot | OpenAI model training | Your call — no effect on being cited today |
| ClaudeBot / Claude-User | Anthropic's index and live fetches for Claude | Yes |
| PerplexityBot / Perplexity-User | Perplexity answers and citations | Yes |
| Googlebot | Google Search AND AI Overviews / AI Mode | Yes — blocking it removes you from Google entirely |
| Google-Extended | Gemini training (not Search) | Your call |
| Bytespider, CCBot | Third-party scrapers and training sets | Block if you like — no visibility cost |

[Every crawler, one page each: what it is and how to control it](https://see-geo.com/bots) · [The platform toggles that block the wrong crawlers](https://see-geo.com/blog/website-platform-ai-visibility-defaults)

## How do I verify the wall is actually open?

Test from outside, as a crawler would — not from your browser, which is exactly the client the wall was built to admit. Run these from any terminal (or an online HTTP tester) and compare the responses:

A healthy answer is a 200 status with your real HTML. A 403, a challenge page, or a response with the vendor's mitigation header means the crawler is still blocked. If your wall verifies bots by IP range rather than user-agent, a spoofed user-agent from your laptop may still be challenged even though the real crawler gets through — in that case the definitive test is the vendor's own bot analytics, or simply re-running the audit and checking the crawler table.

```
# As ChatGPT's search crawler:
curl -sI -A "OAI-SearchBot/1.0" https://yoursite.com/ | head -5
# As Claude's crawler:
curl -sI -A "ClaudeBot/1.0" https://yoursite.com/ | head -5
# As a plain browser, for comparison:
curl -sI -A "Mozilla/5.0" https://yoursite.com/ | head -5
# SeeGeo's own crawler, if you want the audit itself to get through:
curl -sI -A "SeeGeoAudit/1.0" https://yoursite.com/ | head -5
```

[Re-run the free audit — the crawler table is the receipt](https://see-geo.com/)

## Frequently asked questions

### What is the x-datadome header?

It's the response header DataDome adds when it has intervened on a request — typically alongside a 403 and a captcha or JavaScript challenge. If a crawler-user-agent request to your site returns that header, the crawler was blocked; its absence on a 200 response means the request reached your origin.

### Does DataDome distinguish AI assistants from AI scrapers?

Yes, in recent versions of its known-bots settings: AI assistants and search-style crawlers (the ones that cite you) are a separate category from AI data scrapers (training collectors). Allow the first category to stay visible in AI answers; the second is a policy choice with no visibility cost.

### Why does a custom allow rule for a user-agent work if scrapers can fake user-agents?

It's a trade-off. A user-agent allow rule admits anything claiming to be that crawler, so scope it narrowly and prefer DataDome's own verified-bot categories where they exist — they check the crawler's published IP ranges, which a spoofer can't fake. Use the custom rule only for crawlers DataDome hasn't categorized yet.

## Related pages

- [How do I unblock AI crawlers on Cloudflare?](https://see-geo.com/unblock/cloudflare)
- [How do I unblock AI crawlers on Akamai?](https://see-geo.com/unblock/akamai)
- [How do I unblock AI crawlers on Imperva?](https://see-geo.com/unblock/imperva)

---

# How do I unblock AI crawlers on Imperva?
Source: https://see-geo.com/unblock/imperva · Updated 2026-09-02

On Imperva Cloud WAF (the product many still know as Incapsula), AI crawlers are blocked when Bot Access Control blocks their known-bot category, when "challenge suspected bots" is on and the crawler isn't a recognized good bot, or when a security rule matches its user-agent — and the fix is to allow the search-engine and AI-crawler categories, add an allow rule for anything unrecognized, and verify with a crawler-user-agent request. Changes propagate within minutes.

## Why is Imperva blocking AI crawlers on my site?

Imperva classifies clients as humans, good bots (a known-bots list with categories), and bad or suspected bots, and its default policy blocks bad bots and challenges suspected ones with a JavaScript test no crawler can pass. An AI crawler that isn't on the good-bots list, or whose category is set to block, receives Imperva's "Request unsuccessful. Incapsula incident ID" page — and the response carries the x-iinfo header, which is how SeeGeo's audit recognizes an Imperva wall.

Imperva added AI and LLM crawler categories to its known-bots handling as those crawlers became common, so most accounts can now make one category-level decision; before that, each crawler needed its own exception.

## Where does the block live in the Imperva Cloud Security Console?

Paths are current as of mid-2026; Imperva has reorganized the console more than once, so search the settings for these terms if the menus differ.

- Websites → your site → Security → Bot Access Control: the known-bots categories (search engines, monitoring, AI / LLM crawlers…) with Allow or Block per category, plus the "Block bad bots" and "Challenge suspected bots" switches.
- Security → Security Rules (or Policies → Custom Rules): conditions on client type, user-agent, URL or country with an Allow, Block or Challenge action — where blanket "Crawler → Block" rules live.
- Security → Allowlist / Exceptions: IP, URL or user-agent exceptions to the policies above.
- Advanced Bot Protection (if licensed): its own policies per path, which sit on top of Bot Access Control.

## How do I allow AI crawlers on Imperva, step by step?

Save each change; Imperva applies them to its edge within a few minutes with no separate activation step.

- 1. Bot Access Control → known bots: set Search Engines and the AI / LLM crawler category to Allow. If your console shows separate assistant and scraper categories, allow assistants and decide scrapers on policy.
- 2. If "Challenge suspected bots" is on, keep it — but add an exception so recognized crawlers aren't treated as suspected: Security Rules → new rule → client type is Crawler and user-agent contains OAI-SearchBot (repeat or combine for ChatGPT-User, ClaudeBot, PerplexityBot) → action Allow.
- 3. Security Rules: read any rule with action Block or Challenge that matches client type Crawler or a broad user-agent pattern, and exempt the crawlers above.
- 4. If Advanced Bot Protection is enabled, check its per-path policies for the homepage and key pages — they can override the site-level allow.
- 5. Verify with the crawler-user-agent requests below: a 200 with your HTML and no x-iinfo header means the crawler is through.
- 6. Re-run the audit and confirm the crawler table.

## Which AI crawlers should I allow, and which can I keep blocking?

The distinction that matters is search versus training. Search crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Googlebot) are what makes an assistant able to find and cite you; blocking them makes you invisible in AI answers. Training crawlers (GPTBot, Google-Extended, CCBot) feed model training and blocking them costs you nothing in visibility today. Most "Imperva blocks AI bots" settings treat both groups as one, which is exactly why owners who only meant to opt out of training end up invisible.

| Crawler | What it feeds | Allow? |
|---|---|---|
| OAI-SearchBot | ChatGPT search answers and citations | Yes — this is the one that recommends you |
| ChatGPT-User | Live fetches when a user asks ChatGPT about a page | Yes |
| GPTBot | OpenAI model training | Your call — no effect on being cited today |
| ClaudeBot / Claude-User | Anthropic's index and live fetches for Claude | Yes |
| PerplexityBot / Perplexity-User | Perplexity answers and citations | Yes |
| Googlebot | Google Search AND AI Overviews / AI Mode | Yes — blocking it removes you from Google entirely |
| Google-Extended | Gemini training (not Search) | Your call |
| Bytespider, CCBot | Third-party scrapers and training sets | Block if you like — no visibility cost |

[Every crawler, one page each: what it is and how to control it](https://see-geo.com/bots) · [The platform toggles that block the wrong crawlers](https://see-geo.com/blog/website-platform-ai-visibility-defaults)

## How do I verify the wall is actually open?

Test from outside, as a crawler would — not from your browser, which is exactly the client the wall was built to admit. Run these from any terminal (or an online HTTP tester) and compare the responses:

A healthy answer is a 200 status with your real HTML. A 403, a challenge page, or a response with the vendor's mitigation header means the crawler is still blocked. If your wall verifies bots by IP range rather than user-agent, a spoofed user-agent from your laptop may still be challenged even though the real crawler gets through — in that case the definitive test is the vendor's own bot analytics, or simply re-running the audit and checking the crawler table.

```
# As ChatGPT's search crawler:
curl -sI -A "OAI-SearchBot/1.0" https://yoursite.com/ | head -5
# As Claude's crawler:
curl -sI -A "ClaudeBot/1.0" https://yoursite.com/ | head -5
# As a plain browser, for comparison:
curl -sI -A "Mozilla/5.0" https://yoursite.com/ | head -5
# SeeGeo's own crawler, if you want the audit itself to get through:
curl -sI -A "SeeGeoAudit/1.0" https://yoursite.com/ | head -5
```

[Re-run the free audit — the crawler table is the receipt](https://see-geo.com/)

## Frequently asked questions

### What is the x-iinfo header?

It's the response header Imperva's edge attaches to requests it handled, and it appears prominently on its block and challenge pages ("Request unsuccessful. Incapsula incident ID…"). A crawler-user-agent request that comes back with that page was stopped at Imperva's edge and never reached your site.

### Is Imperva Cloud WAF the same thing as Incapsula?

Yes. Incapsula was acquired and rebranded as Imperva Cloud WAF; the challenge pages, incident IDs and some console labels still carry the Incapsula name, which is why SeeGeo's audit reports the wall as "Imperva/Incapsula".

### Can I allow AI crawlers on some pages only?

Yes. Security rules and Advanced Bot Protection policies can be scoped by URL, so you can allow crawlers on your public content pages while keeping stricter handling on login, checkout or API paths — which is usually the right shape: the pages you want cited are exactly the public ones.

## Related pages

- [How do I unblock AI crawlers on Cloudflare?](https://see-geo.com/unblock/cloudflare)
- [How do I unblock AI crawlers on Akamai?](https://see-geo.com/unblock/akamai)
- [How do I unblock AI crawlers on DataDome?](https://see-geo.com/unblock/datadome)

---

# AI visibility tools compared: SeeGeo, Peec AI, Profound
Source: https://see-geo.com/compare · Updated 2026-09-08

Plain comparisons of AI visibility tools by price, buyer and what they measure: SeeGeo vs Peec AI, Profound and Otterly, alternatives, 2026 prices.

## Pages

- [SeeGeo vs Peec AI](https://see-geo.com/compare/seegeo-vs-peec-ai) — Peec AI tracks AI-engine mentions for marketing teams; SeeGeo audits your site and tracks mentions for small businesses. Compared plainly, prices dated.
- [SeeGeo vs Profound](https://see-geo.com/compare/seegeo-vs-profound) — Profound is enterprise AI-search analytics from $99/mo; SeeGeo is a free site audit plus tracking for small businesses. What each measures and costs.
- [SeeGeo vs Otterly](https://see-geo.com/compare/seegeo-vs-otterly) — Otterly monitors AI search mentions from $29/mo for 15 prompts; SeeGeo audits your site free and tracks from $49/mo. Which fits a small business.
- [Peec AI alternatives](https://see-geo.com/compare/peec-ai-alternatives) — Peec AI alternatives compared by price, buyer and what they measure: Profound, Otterly, Semrush's AI toolkit, Rankscale and SeeGeo, with dated figures.
- [How much do AI visibility tracking tools cost?](https://see-geo.com/compare/ai-visibility-tools-pricing) — AI visibility tool prices in 2026 from vendors' own pages: Rankscale $20, Otterly $29, Semrush $99, Profound $99–$399/mo, SeeGeo free audit. Tier by tier.

---

# SeeGeo vs Peec AI
Source: https://see-geo.com/compare/seegeo-vs-peec-ai · Updated 2026-09-08

Peec AI is an AI-search analytics platform built for marketing teams and SEO agencies that track many brands and prompts; SeeGeo is a free site audit plus mention tracking built for one small business that wants to know what to fix. If you run an agency, look at Peec; if you run a business, start with the free audit.

## What does each tool actually do?

Peec AI measures how AI assistants talk about brands: which prompts mention you, how often, in what position against competitors, with which sources cited — presented as dashboards for teams that manage that visibility as a channel. It answers 'how are we doing in AI search this week?'

SeeGeo starts one step earlier. Its audit crawls your site the way an AI crawler does and grades whether an assistant could read, understand and cite you at all — crawler access, structured data, extractability, entity clarity — then lists the fixes. Tracking of mentions and citations on ChatGPT, Claude and Gemini is the second half, so you can see whether the fixes change the outcome.

## How do SeeGeo and Peec AI compare?

Prices read from each vendor's public pricing page on 8 September 2026; check the vendor's site before deciding — this category changes prices often.

|  | Price | Built for | What you get |
|---|---|---|---|
| SeeGeo | Free audit (no account for the score); tracking from $49/mo; Done-for-you $750/mo | Small businesses and startups; agencies via Done-for-you | Site audit + scheduled mention/citation tracking on ChatGPT, Claude, Gemini |
| Peec AI | Published on peec.ai/pricing; tiered by prompts and brands, aimed at teams | Marketing teams and SEO agencies | Brand-mention analytics across AI engines, competitor share, source analysis, team dashboards |

## Who should choose which?

Choose Peec AI if you are an agency or an in-house team that already knows your site is readable and needs to report AI share of voice across many prompts, markets and brands.

Choose SeeGeo if you are a small business owner who does not yet know why ChatGPT recommends a competitor, wants a graded list of fixes in 20 seconds, and would rather have someone implement them than manage a dashboard. The two are not really rivals: SeeGeo is the fix-first tool at the small end of the market; Peec is the measurement platform at the team end.

[Run the free SeeGeo audit](https://see-geo.com/geo-audit) · [Done-for-you: we implement the fixes](https://see-geo.com/done-for-you)

## Frequently asked questions

### Is SeeGeo a Peec AI alternative?

For a small business, yes: SeeGeo audits your site and tracks whether ChatGPT, Claude and Gemini mention you, starting free. For an agency tracking many brands, Peec AI's team dashboards are the closer fit.

### Does Peec AI audit your website?

Peec AI's focus is measuring mentions and citations across AI engines rather than crawling your site for fixes; SeeGeo does the crawl-and-grade step and lists what to change, then tracks the result.

### Which is cheaper, SeeGeo or Peec AI?

SeeGeo's audit is free and tracking starts at $49 a month; Peec AI publishes team-oriented tiers on its pricing page. Compare on what you need: fixes and a single-business view, or multi-brand analytics.

## Related pages

- [SeeGeo vs Profound](https://see-geo.com/compare/seegeo-vs-profound)
- [Peec AI alternatives](https://see-geo.com/compare/peec-ai-alternatives)
- [How much do AI visibility tracking tools cost?](https://see-geo.com/compare/ai-visibility-tools-pricing)

---

# SeeGeo vs Profound
Source: https://see-geo.com/compare/seegeo-vs-profound · Updated 2026-09-08

Profound is an enterprise-grade AI visibility platform — answer-engine analytics, agent traffic, conversation data — with plans from $99 a month to custom enterprise contracts; SeeGeo is a free audit that tells a small business what to fix, plus tracking from $49 a month. Different buyers, different price points, and a fair comparison says so.

## What does each tool actually do?

Profound tracks how AI answer engines describe and cite brands at scale: prompt volumes, share of voice, citation sources, and the traffic AI agents send, packaged for enterprise marketing and SEO teams. It is measurement infrastructure for companies with a channel to manage.

SeeGeo audits a website the way an AI crawler reads it and grades six categories — access, technical, structured data, extractability, off-site, entity — then tracks mentions and citations on ChatGPT, Claude and Gemini so the owner can see the fixes land. It is built for a business with one site and no analyst.

## How do SeeGeo and Profound compare on price and fit?

Prices read from each vendor's public pricing page on 8 September 2026; check the vendor's site before deciding — this category changes prices often.

|  | Price | Built for | What you get |
|---|---|---|---|
| SeeGeo | Free audit (no account for the score); tracking from $49/mo; Done-for-you $750/mo | Small businesses and startups; agencies via Done-for-you | Site audit + scheduled mention/citation tracking on ChatGPT, Claude, Gemini |
| Profound | $99/mo and $399/mo self-serve tiers (billed yearly, two months free), Enterprise tailored | Enterprise and mid-market marketing teams | Answer-engine analytics, share of voice, citation and agent-traffic data, team features |

## Who should choose which?

Choose Profound if you are an enterprise or a large marketing team that needs AI-search analytics as a reporting layer with the depth and support that implies.

Choose SeeGeo if you are a small business or startup that first needs to know whether AI can read your site at all, wants the fixes listed in plain language, and may want them implemented for you. Many small businesses get their biggest gain from the free audit alone — a blocked crawler or a missing About page — before any tracking is worth paying for.

[Run the free SeeGeo audit](https://see-geo.com/geo-audit) · [See SeeGeo pricing](https://see-geo.com/pricing)

## Frequently asked questions

### Is Profound worth it for a small business?

Usually not yet: Profound is priced and designed for teams managing AI visibility as a channel. A small business gets more from a free audit that shows what to fix, and from tracking a handful of the questions its customers actually ask.

### What does Profound cost?

As of 8 September 2026 Profound lists $99 and $399 per month self-serve tiers billed yearly, plus tailored Enterprise pricing; check tryprofound.com/pricing for current figures.

### Does SeeGeo track the same AI engines as Profound?

SeeGeo tracks mentions and citations on ChatGPT, Claude and Gemini with web grounding on; Profound covers additional answer engines and agent traffic at enterprise scale. For a small business the three SeeGeo tracks are where the recommendations happen.

## Related pages

- [SeeGeo vs Peec AI](https://see-geo.com/compare/seegeo-vs-peec-ai)
- [SeeGeo vs Otterly](https://see-geo.com/compare/seegeo-vs-otterly)
- [How much do AI visibility tracking tools cost?](https://see-geo.com/compare/ai-visibility-tools-pricing)

---

# SeeGeo vs Otterly
Source: https://see-geo.com/compare/seegeo-vs-otterly · Updated 2026-09-08

Otterly is an AI search monitoring tool with self-serve plans from $29 a month for 15 prompts, aimed at marketers and agencies; SeeGeo is a free site audit that tells you what to fix, with tracking from $49 a month. Otterly measures; SeeGeo diagnoses first, then measures.

## What does each tool actually do?

Otterly monitors search prompts across AI engines and reports brand mentions, link citations and sentiment over time, with plan sizes set by how many prompts you track — 15 on the Lite plan, 100 on Standard, 400 on Premium. It is a monitoring product: it tells you what the answers say.

SeeGeo's centre of gravity is the audit: a crawl of your site as an AI crawler sees it, six graded categories, and a prioritised list of fixes in plain language. Tracking of your customers' questions on ChatGPT, Claude and Gemini comes with the paid plans so you can watch the fixes land; Done-for-you implements them.

## How do SeeGeo and Otterly compare?

Prices read from each vendor's public pricing page on 8 September 2026; check the vendor's site before deciding — this category changes prices often.

|  | Price | Built for | What you get |
|---|---|---|---|
| SeeGeo | Free audit (no account for the score); tracking from $49/mo; Done-for-you $750/mo | Small businesses and startups; agencies via Done-for-you | Site audit + scheduled mention/citation tracking on ChatGPT, Claude, Gemini |
| Otterly | Lite $29/mo (15 prompts), Standard $189/mo (100), Premium $489/mo (400), Enterprise from $1,000/mo | Marketers, small teams and agencies | Prompt monitoring across AI engines, mention and citation reports, sentiment, alerts |

## Who should choose which?

Choose Otterly if your site is already readable and you want a monitoring dashboard for a set of prompts at a low entry price.

Choose SeeGeo if you do not yet know why you are absent from AI answers: the audit finds the blocked crawler, the missing identity statement or the unquotable pages first, and the tracking then shows whether fixing them changed the answers. If you want neither dashboards nor homework, Done-for-you does the work.

[Run the free SeeGeo audit](https://see-geo.com/geo-audit) · [Done-for-you](https://see-geo.com/done-for-you)

## Frequently asked questions

### Is Otterly cheaper than SeeGeo?

Otterly's Lite plan is $29 a month for 15 prompts and SeeGeo's tracking starts at $49 a month, but SeeGeo's audit — the part that tells you what to fix — is free with no account needed for the score. Compare on what you need, not the first tier.

### Does Otterly tell you what to fix on your website?

Otterly reports what AI answers say about you; diagnosing why — a firewall blocking ClaudeBot, no About page, text that only renders with JavaScript — is what SeeGeo's audit is built to do.

### Can I use SeeGeo and Otterly together?

Yes. Some teams use SeeGeo's audit to fix the site and Otterly or another monitor for large prompt sets; SeeGeo's own tracking covers the handful of questions a small business actually needs to watch.

## Related pages

- [SeeGeo vs Peec AI](https://see-geo.com/compare/seegeo-vs-peec-ai)
- [Peec AI alternatives](https://see-geo.com/compare/peec-ai-alternatives)
- [How much do AI visibility tracking tools cost?](https://see-geo.com/compare/ai-visibility-tools-pricing)

---

# Peec AI alternatives
Source: https://see-geo.com/compare/peec-ai-alternatives · Updated 2026-09-08

The main alternatives to Peec AI for AI visibility tracking are Profound (enterprise analytics from $99/mo), Otterly (monitoring from $29/mo), Semrush's AI Visibility Toolkit ($99/mo), Rankscale (credits from $20/mo) and SeeGeo (free site audit, tracking from $49/mo). Which one fits depends less on features than on who you are: an agency, an enterprise team, or a business with one site.

## What are the alternatives to Peec AI?

Prices read from each vendor's public pricing page on 8 September 2026; check the vendor's site before deciding — this category changes prices often.

| Tool | Price | Built for | Core job |
|---|---|---|---|
| Peec AI | Published on peec.ai/pricing; team tiers by prompts and brands | Marketing teams and SEO agencies | AI-search brand analytics and competitor share |
| Profound | $99/mo and $399/mo (yearly), Enterprise tailored | Enterprise marketing teams | Answer-engine analytics, citations, agent traffic |
| Otterly | $29/mo (15 prompts) to $489/mo (400); Enterprise from $1,000/mo | Marketers and agencies | Prompt monitoring, mentions, sentiment |
| Semrush AI Visibility Toolkit | $99/mo per domain; +$60/mo per 50 extra prompts | Existing Semrush users | AI mention and prompt tracking inside Semrush |
| Rankscale | From $20/mo; Pro $99/mo, then $385 and $780/mo by credits | SEO professionals and agencies | Credit-based AI visibility queries and reports |
| SeeGeo | Free audit (no account for the score); tracking from $49/mo; Done-for-you $750/mo | Small businesses and startups; agencies via Done-for-you | Site audit + scheduled mention/citation tracking on ChatGPT, Claude, Gemini |

## How do you choose between them?

Ask three questions. Do you need to fix a site or measure a channel? Every tool above except SeeGeo is a measurement product; SeeGeo starts with a graded audit and a fix list, then measures. How many prompts and brands? Agencies need hundreds and multi-brand views — Peec, Profound, Otterly's higher tiers. One business needs the ten questions its customers ask. Who will act on the report? If the answer is 'nobody, we are busy', a done-for-you service beats a dashboard.

## Where does SeeGeo sit among them?

At the small-business end, and honestly so. SeeGeo's audit is free, deterministic and versioned; it grades six categories against a published study of 107 real small-business sites and lists what to change. Tracking covers ChatGPT, Claude and Gemini with web grounding. It does not try to be an enterprise analytics suite; it tries to be the thing a business owner runs first and the service that does the work if they'd rather not.

[Run the free SeeGeo audit](https://see-geo.com/geo-audit) · [How the SeeGeo score is calculated](https://see-geo.com/methodology)

## Frequently asked questions

### What is the cheapest alternative to Peec AI?

For a free start, SeeGeo's audit costs nothing and needs no account for the score. Among paid monitors, Rankscale starts around $20 a month and Otterly at $29 a month for 15 prompts, as published on 8 September 2026.

### Is Semrush's AI toolkit a Peec AI alternative?

Yes for teams already on Semrush: the AI Visibility Toolkit is $99 a month per domain and tracks AI mentions and prompts inside the tool you already use, with extra prompts sold in blocks of 50 for $60 a month.

### Which Peec AI alternative is best for a small business?

One that tells you what to fix before charging you to watch the results. SeeGeo's free audit does that in about 20 seconds; if you then want monitoring, its tracking or Otterly's Lite plan are the lowest-cost options.

## Related pages

- [SeeGeo vs Peec AI](https://see-geo.com/compare/seegeo-vs-peec-ai)
- [SeeGeo vs Profound](https://see-geo.com/compare/seegeo-vs-profound)
- [SeeGeo vs Otterly](https://see-geo.com/compare/seegeo-vs-otterly)
- [How much do AI visibility tracking tools cost?](https://see-geo.com/compare/ai-visibility-tools-pricing)

---

# How much do AI visibility tracking tools cost?
Source: https://see-geo.com/compare/ai-visibility-tools-pricing · Updated 2026-09-08

As of 8 September 2026, AI visibility tools range from free (SeeGeo's site audit) through self-serve monitoring at $20–$99 a month (Rankscale, Otterly, Semrush's AI toolkit, Profound's entry tier) to $385–$1,000+ a month for agency and enterprise plans, with done-for-you implementation at $750 a month. The price mostly tracks how many prompts and brands you monitor, not how much your site improves.

## What do AI visibility tools cost, tier by tier?

Prices read from each vendor's public pricing page on 8 September 2026; check the vendor's site before deciding — this category changes prices often.

| Tool | Entry tier | Mid tier | Top tier |
|---|---|---|---|
| SeeGeo | Free audit; Starter $49/mo | Growth $149/mo | Done-for-you $750/mo or $7,500/yr |
| Rankscale | From $20/mo | Pro $99/mo | $385 and $780/mo by credits |
| Otterly | Lite $29/mo (15 prompts) | Standard $189/mo (100 prompts) | Premium $489/mo; Enterprise from $1,000/mo |
| Semrush AI Visibility Toolkit | $99/mo per domain | +$60/mo per 50 prompts; +$99 per extra domain | Team licences at $99 per user |
| Profound | $99/mo (yearly) | $399/mo (yearly) | Enterprise, tailored |
| Peec AI | Published on peec.ai/pricing | Tiered by prompts and brands | Team and agency plans |

## What are you actually paying for?

Prompts and surfaces. Every monitoring tool prices by how many questions it asks the AI engines on your behalf and how often, because each grounded query costs the vendor money — a fact SeeGeo's own cost meter makes visible. More prompts, more brands and more engines push you up the tiers; the site itself is not what is being measured.

That is why the free audit exists as a separate step. Fixing a blocked crawler or a page with no readable text costs nothing to detect and changes every answer at once; monitoring tells you it happened.

## How much should a small business spend?

Start at zero: run a free audit and fix what it finds. Then track the five to ten questions your customers ask, which every tool's entry tier covers — $20 to $49 a month. Move up only when you are tracking competitors across markets or many brands. If nobody in the business will implement fixes, a done-for-you service at a few hundred dollars a month replaces both the tool and the time.

[SeeGeo pricing](https://see-geo.com/pricing) · [Run the free audit](https://see-geo.com/geo-audit)

## Frequently asked questions

### How much do AI visibility tracking tools cost?

From free to over $1,000 a month. Self-serve monitoring starts around $20–$29 a month (Rankscale, Otterly), Semrush's AI toolkit is $99 a month, Profound $99–$399 a month, and enterprise plans are tailored; SeeGeo's site audit is free with tracking from $49 a month.

### Is there a free AI visibility checker?

SeeGeo's audit is free and shows your grade, category scores and crawler table without an account; its crawler check at /ai-crawler-check tests AI bot access alone. Several monitors offer trials, but ongoing prompt tracking is paid everywhere.

### Why do AI visibility tools charge per prompt?

Because each tracked question is a paid, web-grounded query to ChatGPT, Gemini or another engine, repeated on a schedule. Grounding is the bulk of the cost, so plans scale with prompts and frequency rather than with features.

## Related pages

- [Peec AI alternatives](https://see-geo.com/compare/peec-ai-alternatives)
- [SeeGeo vs Otterly](https://see-geo.com/compare/seegeo-vs-otterly)
- [SeeGeo vs Profound](https://see-geo.com/compare/seegeo-vs-profound)
