We Asked 6 AI Engines the Same Buyer Question. We Got 6 Different Markets.
We ran one buyer question — which firm to shortlist for a GTM + RevOps + AI-search operating partner — through six AI engines. Google AI and Perplexity answered with venture capital firms. ChatGPT, Copilot, Grok and Claude answered with agencies. No provider showed up in all six. The failure isn't ranking — it's legibility.
By Kunal Achintya Reddy · 3 min read · 3 September 2026

We Asked 6 AI Engines the Same Buyer Question. We Got 6 Different Markets.
Not 6 different answers. 6 different interpretations of what we were even asking for.
The prompt
A 50-person, VC-backed B2B SaaS company (Europe/India/US) needs one senior operating partner to own GTM strategy, RevOps, and AI-search visibility. Which firms should we shortlist, and why?
We ran this through six engines: ChatGPT Search, Claude, Google AI, Microsoft Copilot, Perplexity, and Grok.
What happened
| Engine | Read "operating partner" as... |
|---|---|
| Google AI | A venture capital role — recommended investment firms |
| Perplexity | A venture capital role — recommended investment firms |
| ChatGPT | An agency / consultancy / fractional executive |
| Copilot | An agency / consultancy / fractional executive |
| Grok | An agency / consultancy / fractional executive |
| Claude | An agency / consultancy / fractional executive |
The six answers split into two camps before they ever got to recommending a specific firm. Google AI and Perplexity treated "operating partner" as a venture capital title and pointed us to investment platforms. ChatGPT, Copilot, Grok, and Claude treated it as a services role and pointed us to agencies, consultancies, and fractional executives.
Claude went a step further and flagged something the other five didn't surface explicitly: GTM plus RevOps is a coherent single mandate, but AI-search visibility is usually sold as a separate specialist engagement — meaning the exact combined offer we described barely registers as an existing category yet.
Several answers also pointed to automation tooling or CRM features as "evidence" of AI-search expertise. Those are not the same job, and conflating them is itself a legibility failure — the model reaching for the nearest adjacent thing it could confidently classify.
No single provider showed up across all six answers. The only pattern shared by all six was the fork itself: investor platforms on one side, service firms on the other.
This isn't a ranking problem — it's a legibility problem
Most AEO/GEO work assumes the hard part is earning a citation once a model has decided what you are. This field study suggests the harder, earlier problem is getting six different systems to agree on what you are in the first place.
Before a brand can be cited or recommended, the model has to correctly classify the company. If six systems can't converge on that from one clear, specific prompt, then a brand visibility score built on a single platform — "we rank in ChatGPT" — is measuring the wrong layer of the problem. You can be perfectly optimized for citation inside a category the model has assigned you to incorrectly.
Citing a source didn't fix it either
Several answers leaned on provider-owned marketing pages or thin directory data as their cited evidence. A citation tells you where a model pulled a claim from — it does not tell you the claim is accurate, or that the source is independent of the thing being described. An answer can be well-cited and still be classifying you into the wrong market.
What survives if the AI hype cycle corrects
If the current wave of AI-search hype corrects, the fragile version of AEO — screenshot-driven "we rank in ChatGPT" claims — goes first. What survives is the boring, durable work:
- Can a crawler actually access your information?
- Does the independent web (not your own marketing pages) accurately describe what you do?
- Does that description hold up consistently across multiple systems — not one?
That third question is the one this field study was built to test, and it's the one most single-platform visibility tracking skips entirely.
Methodology note
This was a single prompt run once per engine, read for category classification rather than exhaustive citation counting. It's a snapshot, not a longitudinal study — the value is in the disagreement itself, not in ranking which engine "got it right." Treat it as a prompt for auditing your own category legibility across engines, not as a benchmark to optimize a single score against.
