← Blog
10 September 2026 ⌘ 17 min read
Blog 10 September 2026

The Model Already Has a Favorite

Why ranking #1 for 'best help desk software' doesn't get your SaaS mentioned in ChatGPT, and why the industry selling you the fix doesn't fully understand the thing it's selling.

The Model Already Has a Favorite

Let’s call her Jordan. She runs marketing for a mid-market help desk platform, the kind of company that spent two years grinding to the #1 organic spot for “best help desk software for small teams.” Real position one, verified in Search Console, not a screenshot cropped to hide the ads above it. Then one afternoon she opened ChatGPT, typed nearly the exact question her prospects type, and watched it recommend three competitors by name. None of them outranked her that week on Google. Her product didn’t come up at all, not even as a footnote link at the bottom of the answer.

She did what any marketer under quarterly pressure would do. She asked her SEO agency why.

“It’ll take time. Give it another quarter.”

She gave it another quarter. Nothing moved. She asked again.

“We need more mentions in third-party publications. Once you’re cited on a few more comparison sites, the models will pick it up.”

She paid for the comparison-site placements. Nothing moved. She asked a third time.

“Honestly, we probably need more budget to optimize the technical side further.”

That’s the point where she called me, not because I had a magic answer, but because she’d noticed the answers kept changing shape while the underlying problem didn’t move an inch. Ask “why” three times in a row to most people selling AI-visibility services right now, and you’ll watch them talk themselves into a corner, because it’s genuinely one domain almost nobody understands yet, least of all the people billing you monthly for understanding it.

“The map is not the territory,” people say, usually misquoting the general semanticist Alfred Korzybski, who actually wrote something drier and more technical about the structural relationship between an abstraction and what it describes.[1] I’m opening with the popular, imprecise version anyway, because it’s the closest thing I have to a working description of what’s happening between a website and the answer ChatGPT gives about it. The website is the territory. The model’s sense of a brand is a map drawn mostly before it ever visited.

Here’s the thesis, stated once, plainly, so the rest of this piece can argue for it properly:

Optimizing content and links to get mentioned by name in AI answers is a losing game for any brand trying to displace a competitor who already owns the category, because the decision about who gets mentioned is largely made before the model searches anything, and no amount of content velocity you can produce out-scales the training data and live web volume the model is weighing you against.

The three answers, and why they’re all half-true

Go back to Jordan’s agency for a second, because those three answers deserve more credit than “it’ll take time” implies, and less credit than “we need more budget” implies.

It will take time is true in the sense that entity recognition, the process by which a model comes to associate a brand name reliably with a category, does build up gradually across new training runs and repeated web mentions. It is misleading in that “time” is standing in for “we don’t actually know how much time, or whether more time alone fixes it,” which is a very different sentence to sell a retainer on.

We need more third-party mentions is true in that independently published, citation-worthy coverage genuinely does feed both search rankings and future training corpora. It is incomplete in that a mention existing somewhere on the internet and a mention actually shaping what a model retrieves or recalls are two different events, and most agencies report the first one because it’s the only one they can screenshot.

We need more budget to optimize further is the answer that should worry you most, not because more resourcing never helps, but because “optimize further” rarely comes with a specific mechanism attached. Optimize what, exactly, against what measured baseline, expecting what specific change in model behavior? When the next sentence stays vague, you’re not funding a fix. You’re funding continued uncertainty with better production values.

None of this is dishonesty. It’s the sound people make when they’re describing a system from the outside, guessing at the shape of something they can’t open.

Two processes wearing one name

“AI search” gets talked about as one mechanism. It’s actually two, running at different times, and the confusion between them is where most of the wasted budget lives.

The first is what happened while the model was trained. Somewhere in a training run, a brand name, or its absence, got folded into the model’s parameters alongside every other brand name that showed up often enough, in enough authoritative-looking contexts, to leave a statistical trace. This is the model’s memory, and it’s expensive and slow to change. It doesn’t update when you publish a blog post next week. It updates, if it updates at all, the next time a lab trains a new model on a new data snapshot, on a schedule you don’t control and can’t see.

The second is what happens the moment someone hits enter on a question. Many AI answers today lean on more than memory. The model breaks a question into a handful of its own background searches, called fan-out queries, runs them against a live index, reads what comes back, and writes an answer from whatever survives. Chris Long’s team at Nectiv, comparing roughly 4,000 prompts against their own 2025 baseline, found ChatGPT’s average fan-out queries per prompt jumped from 2.17 to 7.61 in a year, with the longest observed chain going from 4 searches to 29, and software queries running the highest fan-out volume of any category tested, 10.7 per prompt.[2] The site: operator alone now shows up in 64% of all fan-out queries.

Here’s the part that should reframe how you think about the whole problem. Suganthan Mohanadasan tracked this directly through OpenAI’s own API and found the model’s own first search query, before it retrieves a single page, often already contains specific brand names the user never typed.[3] Ask about the best AI note-taking app with no product named, and the background search ChatGPT runs on its own might already read something close to a shortlist: Granola, Notion AI, Otter, Fireflies, Fathom, Mem, Limitless, before a single page gets fetched. Across the conversations he tested, 21 of 27 first queries already contained brand names nobody typed. And the citation gap that follows is the real finding, more than anything about the site: operator. Brands the model names itself in that first query reach the final answer 68.9% of the time. Brands that only turn up later, retrieved but never named by the model itself, reach it 2.1% of the time.

That’s a 33-times gap, and it isn’t a ranking gap. It’s a shortlist gap. The model decided who was in the running before it went looking for anyone.

Ranking #1 was never the shortlist

If ranking #1 doesn’t buy a spot on that shortlist, what does correlate with getting cited? Ahrefs ran the broadest version of this test: 15,000 long-tail queries across ChatGPT, Gemini, Copilot, and Perplexity.[4] Only 12% of the URLs those four assistants cited ranked in Google’s top 10 for the original prompt. 80% of citations didn’t rank anywhere in Google for that query at all. Perplexity leaned closest to traditional search, at 28.6% overlap; ChatGPT and Gemini sat closer to 8%. Google’s own AI Overviews, for contrast, pulled 76% of their citations from top-10 pages, which tells you the “does ranking matter” answer depends enormously on which AI surface you’re asking about.

A narrower, buying-intent-specific study from Grow & Convert tested 100 prompts like “best help desk software for small teams,” the exact category Jordan competes in, and checked ChatGPT’s citations against real Google and Bing rankings for those queries.[5] Roughly 40% of citations came from a page ranking anywhere in the first 10 pages of either engine. The rest didn’t. But when the researchers checked correlation at the domain level instead of the exact URL, it jumped from 27% to about 50%. The specific page ChatGPT quotes is often not a brand’s best-ranking page for that query. The domain it’s pulled from usually ranks well for something adjacent. As the authors put it, chasing a single piece of content to rank for one AI prompt is “fuzzier and less direct than that.”[6]

Which means the honest, uncomfortable version of Jordan’s situation is this: she wasn’t failing to rank. She was failing to be a brand the model had already decided was worth naming, and no single well-optimized page fixes that, because the model isn’t grading pages one at a time. It’s recalling a category the way you’d recall which three vendors are “the good help desk tools” without re-researching the question every time a colleague asks.

The people actually paying for a fix nobody can fully see

I want to be specific about who gets hurt by the “publish more, build more links, wait three to six months” pitch, because it isn’t an abstraction. It’s marketing leads like Jordan, under quarterly pressure to show an “AI visibility” number moving, and in-house SEOs asked to defend a line item to a CFO who has started asking, not unreasonably, “are we in ChatGPT yet.”

The honest answer, most of the time, is that the model already decided a shortlist for the category, probably from training data assembled before the last dozen blog posts existed, and the fastest-growing service line in the industry right now is selling incremental fixes against a decision that was mostly made somewhere upstream of anything a content calendar can reach.

I don’t think most people selling this are being dishonest. I think most of them are describing a black box from the outside and calling the shape they see a strategy. Reddit is the cleanest illustration of how deceptive that outside view can be. It holds one of the largest citation shares of any single domain across AI-citation trackers, and if that’s the only number anyone looks at, the conclusion is that AI models love Reddit more than anywhere else on the internet. Dejan.ai pulled the underlying retrieval data straight from OpenAI’s own grounding API instead of scraped chat transcripts, and found Reddit appears as a candidate source in 76% of OpenAI’s searches, with 491,024 Reddit pages retrieved across a six-month window. Only 3,012 got cited. A 99.39% rejection rate.[7] Across the same window, Anthropic’s Claude cited Reddit zero times out of 139,601 grounding sources. Reddit doesn’t win because models prefer it. It wins on volume, surviving a filter that discards it almost every single time it’s offered, and even that survival isn’t universal across models.

If a team is only tracking whether their subreddit thread got pulled into a search, the dashboard says the Reddit strategy is working. It’s failing 99 times out of 100, and the rare hit is carrying the whole story.

What a founder told me after she stopped chasing the mention

A client I’ll identify only as R., who runs marketing for a logistics SaaS company, put it to me more bluntly than I would have.[8] Nine months into a retainer explicitly sold as “AI search optimization,” she asked her agency for the actual mechanism connecting their monthly deliverables to a citation appearing anywhere. What came back was a slide about “authoritativeness signals” and a promise that things “compound.” She canceled the retainer the next week.

What she did instead is the part worth repeating. She pulled sales, support, and two long-tenured customers into one room and built a list, stage by stage, of the actual questions people asked before, during, and after buying. Not keywords. Questions a human had said out loud to another human, on a call. Then her team wrote directly and specifically to what they believed the product should say at each stage, opinions included, without checking a ranking tool first.

Six months later she hadn’t tracked a single AI-mention metric. Her demo-to-close rate had moved, because the content answered questions her sales reps used to answer manually, on calls that no longer needed to happen.

The honest concession: this doesn’t mean stop writing

Someone reading this is thinking, correctly, that content still obviously matters, and that argument deserves its full weight before I push back on it. Better content genuinely helps humans find and trust a product. Better technical structure genuinely makes a site easier to crawl, for people and for machines. None of that stops being true.

Where it stops being sufficient is the specific claim being sold: that enough of it, sustained long enough, reliably displaces a brand the model already associates with a category. A single company’s content output cannot out-scale the training corpus and the live web volume the model is comparing that brand against. The competition isn’t one rival’s blog. It’s every mention of every competitor across a scrape of the internet larger than any team can author its way past, weighted by a recall process nobody outside the labs that built it can fully see.[9] That gap is a scale problem wearing a content problem’s clothes.

How I wrote this

I researched the retrieval and citation studies cited above by reading each primary source directly, at the URLs footnoted here, then used Claude to help cross-check the numbers against each other and flag where two studies measuring a similar window disagreed, which happened at least once in the underlying research.[10] The R. anecdote is disguised at the source’s request, drawn from a real engagement, altered enough that the company isn’t identifiable. The opinions, the framing, the specific claim that this is a scale problem rather than a content problem, and every sentence of the argument itself are mine, edited by hand after the draft, not generated and left unread.

Back to Jordan

Jordan’s product still isn’t mentioned in ChatGPT’s answer to her prospects’ exact question. It might get there eventually, or it might not, and I told her honestly that nobody, including me, can currently tell her which. What changed is what she’s doing about it. She stopped asking her agency why the mention hadn’t arrived, and started asking her actual customers what they needed to know before they’d trust a help desk tool enough to migrate their support queue onto it. Her link-building process now asks one extra question before every placement: would this page’s real readers click through and start a trial, or is this just a link. Her Reddit comments stopped being written for a model that might scrape them and started being written for the one person reading them at midnight, trying to decide between three tools before a renewal deadline.

The shortlist for her category might still be set for a while. But the trust she’s building with the people actually reading her, the ones who ask her a follow-up question in the comments instead of just moving on, doesn’t need an invitation from anyone.

If you’re optimizing for a mention, stop, and go build the thing a mention was only ever supposed to be evidence of.


If this was useful, subscribe. I’m writing a short series on what actually moves brand recognition versus what only moves a dashboard, one honest piece at a time.

Tell me where I’m wrong in the comments, in your own words, not a model’s. That request is itself a small test of the argument above: a comment written because someone actually disagrees is worth more to me than ten that only exist to be cited somewhere.

Next up: what it actually took R.’s team to rebuild a content calendar around real sales-call questions instead of keyword volume, with the calendar itself.

A running list of every study cited across this series lives in the Research Library (link to be added once that page exists).

Notes

  1. Korzybski’s actual 1931 formulation is far more technical than the popular paraphrase suggests, concerned with the logical relationship between an abstraction and the structure it describes, not a general warning against confusing models for reality. The popular version survives because it’s useful shorthand, not because it’s an accurate quotation, a small irony given the subject of this essay.

  2. Chris Long, “New Research: ChatGPT Tripled Its Fan-Out Queries,” Nectiv, August 13, 2026. Roughly 4,000 prompts compared against Nectiv’s own 2025 baseline on the same query set. site:, “official,” and “gov” were the three most common words the model added on its own. A separate, larger study from Peec AI, 5 million fan-out queries across ChatGPT, Perplexity, and Grok collected around the same window, reported a lower average fan-out count and different top injected words, “best,” “what,” and “review.” Neither study is wrong; they measured different platforms and counting methods, which is a good reminder that “how you counted” changes a headline number as much as “what the model did.”

  3. Suganthan Mohanadasan, “ChatGPT Already Knows Who’s in the Running Before It Searches,” published August 10, 2026, updated August 17, 2026. Sample: 27 conversations for the first-query test, 57 conversations and 3,554 retrieved pages for the citation analysis, 110 pages cited, a 3.1% citation rate overall. The author flags the sample explicitly as “one account and a few hundred conversations, weighted towards software and AI tools,” and calls the numbers directional rather than universal. I’d want to see this replicated across more accounts and verticals before treating 68.9% and 2.1% as fixed constants, but a 33-times gap is hard to explain away as noise.

  4. Ahrefs, 15,000 long-tail queries across ChatGPT, Gemini, Copilot, and Perplexity, published August 11, 2025. Copilot’s overlap with Bing’s top 10 specifically was 16.6%, roughly in line with its reliance on Bing’s index rather than an independent crawl.

  5. Katelyn Urich, “Only 40% of ChatGPT Sources Come from Google & Bing for Known Fan-Out Queries,” Grow & Convert, March 30, 2026. Worth noting separately: 78% of what got cited still looked like conventional SEO content, listicles, product pages, homepages, pricing pages, regardless of whether the specific page ranked for the fan-out query that surfaced it. The model isn’t abandoning search-ranked content; it’s pulling more broadly from a trusted domain’s pages than the single URL sitting at position one.

  6. Same study as above. Full sentence, trimmed here to keep the quoted portion short: the authors’ point was that AI search is “fuzzier and less direct” than traditional rank tracking, which is their conclusion, not a data point, and worth reading in full rather than taking my summary of it as gospel.

  7. Dejan.ai, “No, AI Doesn’t Prefer Reddit. Search Does.”, published July 21, 2026, drawing on six months of retrieval data through OpenAI’s grounding API rather than scraped chat sessions. I’m citing this over chat-scrape studies specifically because it measures the funnel Reddit strategies actually depend on, offered as a candidate then either cited or discarded, not just the visible end state.

  8. Composite and disguised at the source’s request, drawn from a real engagement; details altered enough that the company isn’t identifiable, kept true enough that the mechanism is accurate. This is the one place in the piece where the claim rests on my characterization rather than a checkable source, and I think that trade is worth flagging plainly rather than dressing the anecdote up as more documented than it is.

  9. This paragraph is my own extrapolation from the cited studies, not a conclusion any single one of them draws directly. Nobody has published a clean measurement of “training corpus scale versus achievable content output” as a ratio. What the studies do show, consistently, is that pre-existing brand association predicts citation far better than recent content volume does, and the scale argument is my read of why that would be true. That inference is mine, not the researchers’.

  10. Nectiv and Peec AI, both measuring ChatGPT fan-out behavior in mid-2026, reported meaningfully different average query counts per prompt, as noted in footnote 2. Neither dataset is wrong. It’s the cleanest reminder I found while researching this that a study’s counting method can move a headline number as much as the model’s actual behavior does.

// end_of_post
← All posts Book Your AI Visibility Audit →