Perplexity's AI recommendations lean on 215,000 machine-generated 'best software' pages
Three sites made 215,128 "best software" pages for AI. Perplexity cites them

An analysis of 7,534 citations from Perplexity's Sonar models across 380 software categories found that 59.8% point to domains ranked worse than #100,000 on Tranco, and 23.4% to domains outside the top million. Three sites under apparent common control—WorldMetrics, WifiTalents, and Gitnux—published 215,128 generated buying guides, none existing before December 2023. These sites, which self-describe as 'Facts & Grounding Pages,' were cited 181 times. Guideflow, a demo vendor's blog, ranked third overall with 194 citations. The study notes limitations: it covers only Perplexity and one day's snapshot.
These pages are addressed, in their titles and descriptions, to the software that reads them.
- xpct
If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated websites when I ask them to search for something. It also doesn't help that the web search tools that OAI and Anthropic have are deeply limiting: can't exclude keywords or domains.
- mstaoru
Well it's not only this, or protection from LLMs training on LLM output. LLMs training on human output is also problematic.
I was traveling to an obscure small town, doing some "research" with LLMs beforehand. Every and each one told me enthusiastically to go to "Foobar square" (name changed) for the "best street food in XYZ town", some added a lot of colorful details.
There was no Foobar square in XYZ town. There was no Foobar square anywhere in the world. There was a SINGLE old Reddit comment, with no upvotes, to a unpopular post in an unpopular subreddit, where someone clearly badly misspelled the name of the square, and said something like "for street food go to Foobar square". Nothing about "the best" even.
It's all a lie.
- Aurornis
I used one of the 12-month free Perplexity offers when they were everywhere. It felt slightly useful at first for simple queries where I didn’t want to go through the top 10 Google results manually. If I was looking for a specific recipe I remembered or a help page or user manual it would usually find it quickly.
Then they started optimizing for speed of responses over quality of results. I can enter a query and see my results appear in a second, but they’re garbage. The links and references it gives frequently don’t match the text right next to them. It feels like someone had a KPI to make responses as fast as possible and they optimized for that above all else.
They added a “Computer” option that’s supposed to do research for you. Half the time I can’t get it to trigger through the UI. Pressing the submit button doesn’t work. When I can get it to trigger, most of those sessions will work for a while and then just stop before an answer comes back.
The only reason I keep using it is to keep observing a company that has been heavily marketed and hyped, which should have had a market leading position for something. Even non-technical people I know who listen to Joe Rogan (where Perlexity is advertising heavily, I’m told) are asking me about it.
Now there are reports of people being billed at the end of their trial period without warning, despite them saying that they will warn before this happens. There are some alarmingly bad customer support screenshots where the customer sup […]
- toddmorey
I do think models currently don't have enough source skepticism.
If you look at agent traces when asked to compare two options to help inform a decision, many of the comparison pages cited in research are often hosted by one of the companies being compared; nearly all are AI-generated AEO plays. Not deeply considering the motive of published information is currently a glitch that can be exploited, but the window will close.
I'm sure model providers will set up some crappy pay for play verification system for "trusted" product information, comparisons, and reviews.
- jpimbert
It's difficult to read more than a few sentences, when this itself is clearly a Claude artifact.
- alangibson
Perplexity is about to learn that Google is an anti-spam company first, search engine second
- rcar1046
"Sharing a nameserver pair is strong circumstantial evidence of a common Cloudflare account rather than proof of ownership"
-when you read one statement that let's you know to believe no other assertions in the article....
- throwaway2037
This is genius. The AI/LLM singularity has arrived, and it is shaped like a snake eating its own tail (ouroboros) [1] (or a pelican riding a bicycle).
[1] https://www.newsbiscuit.com/post/ouroboros-unclear-if-it-s-e...