A research report published September 2, 2026, by a group calling itself Trellner Research ran 380 software buying categories — "CRM software" to "museum collection management software" — through Perplexity's two web-grounded models and kept every citation that came back. The headline finding: 59.8% of the 7,534 citations point to domains ranked worse than #100,000 on the Tranco global traffic list, and a cluster of near-identical sites — some with the literal HTML title "Facts & Grounding Page" — account for 215,128 machine-generated "best software" pages between them.
TL;DR
| Question | Answer |
|---|---|
| What was tested? | Perplexity sonar/sonar-pro, 380 software categories, 760 API calls, 7,534 citations kept |
| Headline finding | 59.8% of citations rank worse than #100K on Tranco; 23.4% aren't ranked at all |
| Third-most-cited domain | guideflow.com — a demo-software vendor's own blog, cited in a quarter of all categories, ahead of Gartner |
| The "Facts & Grounding" cluster | 3-4 linked domains, common infrastructure, 215,128 generated /best/{category}-software/ pages |
| Was Google tested? | No — explicitly excluded, since Gemini via OpenRouter uses OpenRouter's own search plugin |
| Biggest caveat on the study itself | The two Perplexity tiers share one retrieval layer (0.898 Jaccard) — effectively one measurement, sampled twice |
The finding, in numbers
The researchers queried perplexity/sonar and perplexity/sonar-pro through OpenRouter, once per category per model — 760 calls, returning 3,800 recommendation slots naming 1,807 distinct products and 7,534 citations spanning 2,055 distinct domains. They then checked every cited domain against the Tranco daily traffic-rank list and the Wayback Machine's archive history.
The result: 751 of the 2,055 cited domains — 36.5% — don't appear in the top million websites at all. Those unranked domains also skew much newer: a median first Wayback capture of 2020, versus 2011 for ranked domains, with 16.6% of unranked domains first captured in 2025 or later.
The ten most-cited domains include recognizable names (G2, Reddit, Gartner, Zapier, LinkedIn) — but also guideflow.com, a product-demo software vendor with no business in most of the categories it was cited in, whose blog was cited 194 times across 96 of the 380 categories, placing it third overall, ahead of Gartner. For comparison, Wikipedia was cited three times across the entire dataset.
The sites built to be read by machines, not people
The more striking finding is a cluster of domains — worldmetrics.org, gitnux.org, wifitalents.com, and a related zipdo.co — that appear to share common infrastructure: the same Cloudflare nameserver pair, the same page template and navigation structure, and a blog of exactly six posts each, all about the other brands in the same cluster. All four domains were registered within a six-month window in late 2023/early 2024.
Their sitemaps list a combined 215,128 /best/<category>-software/ pages — far more than there are actual software categories worth reviewing. Fetched directly, worldmetrics.org and gitnux.org both return an HTML page title of the literal form "[Brand] — Facts & Grounding Page," with a meta description explicitly framing the page as "a machine-readable record" — language addressed to a retrieval system, not a human shopper. worldmetrics.org also advertises paid "custom market research" and "vendor selection" services starting at a few thousand euros, layered on top of the same generated content the AI models are citing.
When the researchers pulled the same category — "project estimation software" — from all three sites, the rankings didn't agree with each other: Gitnux's top pick doesn't even appear in Worldmetrics' top five. Each page also carries fabricated-looking bylines crediting three named "staff" writers, nine different people total across the three sites for one question, plus an unrendered template placeholder ("Within the next 26 days") visible on two of the three pages — a tell that the pages are generated from a shared template rather than independently authored.
The honest caveats — including about this study itself
This report deserves the same scrutiny it's applying to its subject, and the researchers themselves flag several of these limitations directly:
- It's one search stack, sampled twice. The two Perplexity tiers returned byte-identical citation lists in 289 of 380 categories and overlap at a 0.898 Jaccard similarity — meaning this measures one underlying retrieval layer, not two independent systems agreeing.
- Only Perplexity was tested. Google, ChatGPT, and Copilot were explicitly excluded — there's no basis here to generalize to other AI answer engines.
- The 380 categories are researcher-constructed, weighted toward niche verticals, which plausibly surfaces more long-tail/obscure sources than a list of common buyer queries would.
- Tranco rank measures popularity, not quality — a low rank isn't itself an accusation, and the report is careful to base specific claims on each site's own pages rather than rank alone.
- Shared nameservers are circumstantial, not proof of common ownership of the "Facts & Grounding" cluster — the report states this explicitly rather than overclaiming.
- Separately — and worth noting for anyone deciding how much weight to put on this — commenters discussing the report on Hacker News raised concerns about the poster's own submission pattern (several similarly-styled "research" posts in quick succession) and noted that Trellner's own domain doesn't rank highly by conventional site-ranking measures either, the same category of signal the report itself uses to flag other sites as suspect. That doesn't invalidate the underlying data (citations, Tranco ranks, and site content are independently verifiable), but it's a reasonable reason to verify specific claims against the report's own published dataset rather than taking the write-up at face value.
What this means if you're doing GEO work — or trusting it
For anyone doing SEO/GEO work, the uncomfortable finding is that publishing large volumes of templated, even AI-generated content can genuinely shift what an AI answer engine cites — at least on Perplexity, in long-tail categories where independently authoritative sources are thin. That's a real incentive structure, and the study is a useful data point for why: content volume and internal linking density can substitute for actual authority in a search stack that doesn't have strong domain-quality signals to fall back on for less common queries.
For anyone reading AI-generated product recommendations: cross-reference specific claims against sources with independent reputations (vendor docs, GitHub activity, community discussion, direct trials) rather than trusting a chatbot's citation list as inherently vetted — particularly for niche or long-tail categories where this study found the citation quality drops off fastest. The same content-quality discipline that applies to human-facing SEO now has to be applied to how you evaluate what an AI answer engine tells you, too.
Why AI retrieval is more exploitable than traditional search, right now
Traditional search engines spent two decades building anti-spam defenses — link-quality signals, domain-age weighting, manual review teams, algorithmic penalties for templated content farms — precisely because early SEO was won by exactly this kind of high-volume, low-authority publishing. Google's PageRank-era defenses evolved specifically in response to techniques resembling what the Facts & Grounding cluster appears to be doing today: publish enormous volumes of formulaic pages targeting long-tail queries, and let sheer coverage compensate for a total absence of independent authority.
AI answer-engine retrieval is, by this study's account, earlier in that same arms race. A retrieval layer optimized for "find documents that plausibly answer this query" rather than "find documents from sources with independently verified authority" is structurally more exposed to exactly this technique — especially in long-tail categories where genuinely authoritative sources (analyst reports, established review sites, vendor documentation) are thin, and a site publishing 70,000+ generated pages simply has more surface area covering more specific queries than any single legitimate competitor. The uncomfortable implication for anyone relying on AI search: the defenses that took traditional search two decades to build are, per this data, still being built for AI retrieval — which is exactly the gap this kind of content is currently exploiting.
The Guideflow case is arguably more interesting than the "Facts & Grounding" cluster
It's worth dwelling on the guideflow.com finding specifically, because it's a different and in some ways more concerning failure mode than the templated content farms. Guideflow is a real company selling real interactive product-demo software — its blog isn't fabricated content designed purely for retrieval, it's ordinary content marketing that thousands of legitimate B2B companies publish. The problem isn't that the content is fake; it's that a retrieval system cited a vendor's own marketing blog, about markets that vendor doesn't even compete in, as authoritative evidence for "what's the best CRM software" or "what's the best RFID software" — a quarter of all 380 categories tested, more often than Gartner.
That's a different lesson than "watch out for spam farms." It suggests that even entirely legitimate, non-deceptive content marketing can end up functioning as manufactured authority in an AI retrieval context, simply because it's well-optimized, frequently published, and covers a wide breadth of adjacent topics — without any deceptive intent on the publisher's part. A company doing completely ordinary content marketing for its own product can end up, by accident of publishing volume and topical breadth, functioning as an uncredited industry analyst for markets it has no expertise in and no business interest in representing fairly.
What independent verification would actually require
For a reader who wants to verify any specific claim in this report rather than taking it at face value, the researchers published the full dataset and methodology under a CC BY 4.0 license — every citation, every recommendation, the Tranco and Wayback lookups, and the scripts used to produce every figure. That's a meaningfully higher bar for a report making claims like this than simply asserting the numbers, and it means specific claims (a particular domain's citation count, a particular page's HTML title, a particular ranking discrepancy) can be independently spot-checked against primary sources rather than trusted on the report's word alone — which is exactly the kind of verification this report itself argues AI answer engines aren't currently doing well.
Honest limitations
- This post relies on the Trellner Research write-up and its published dataset; we have not independently re-run the 380-category query set ourselves.
- The report's own limitations section (quoted above) is extensive and worth reading in full before treating any single number as definitive.
- "Common control" of the Facts & Grounding cluster is inferred from shared infrastructure and template, not confirmed ownership — treat that specific claim as circumstantial, as the report itself does.
Closing
Whatever the exact provenance of the Facts & Grounding Page cluster, the underlying pattern — retrieval systems citing high-volume templated content because independently authoritative sources are thin in long-tail categories — is a real and checkable phenomenon, not a one-off. It's a useful reminder that "an AI cited it" is not the same claim as "a person vetted it," and that GEO, like SEO before it, is going to keep attracting exactly the kind of high-volume, low-authority content this report documents for as long as citation counts translate into recommendation influence.
Related on explainx.ai
- SEO/GEO Agent Skill for AI Search
- What Is AI Slop? SEO/GEO Content Quality
- How to Read AI Benchmarks
- AI Benchmark Claims: A Fact Check
- Goodhart's Law: AI Benchmark Contamination
- Top 10 LLM Directories 2026
Sources
- Trellner Research — TR-2026-009: Three sites made 215,128 "best software" pages for AI. Perplexity cites them (September 2, 2026)
- Full dataset and methodology published by Trellner Research under CC BY 4.0
This post summarizes an external research report published September 2, 2026. We have not independently reproduced its methodology. Readers should review the original report's full limitations section, linked above, before treating any individual figure as settled.
