Live data from Hacker News

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

trellner.com

81–90 of 265 posts

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#81
post #74
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

A bit different, but one thing I’ve seen is models repackaging Reddit slop. Like, it will do a search, find a Reddit thread somewhat related where someone in a comment casually mentioned incorrect information that any human would have dismissed. The model takes that as granted, but expands on it and present it as a well established fact, presented in a very plausible fashion. In general I don’t find models to be good…

i actively assume it is worse since, for example, spez signed a 60 million dollar deal to give Google access to the firehose. so then if you have niche, highly engaged subreddits infested by AI bots creating posts, then commenting on posts, then being trained on that content... you have Ouroburos eating its own poop, and models have less then zero incentive to evaluate the quality of a source, especially if they are the source

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#83
post #74
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

A bit different, but one thing I’ve seen is models repackaging Reddit slop. Like, it will do a search, find a Reddit thread somewhat related where someone in a comment casually mentioned incorrect information that any human would have dismissed. The model takes that as granted, but expands on it and present it as a well established fact, presented in a very plausible fashion. In general I don’t find models to be good…

Newspapers have been doing this for a long time, notably the Metro in London.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#84

I do think models currently don't have enough source skepticism. If you look at agent traces when asked to compare two options to help inform a decision, many of the comparison pages cited in research are often hosted by one of the companies being compared; nearly all are AI-generated AEO plays. Not deeply considering the motive of published information is currently a glitch that can be exploited, but the window will…

Yes 100%. I see this all the time. You’re asking about a product and the LLM will cite a source from a competitor where the competitor will review the source and list a few positives about the product but lots of negatives. Then the LLM uses them in the response. So cheeky

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#85

Earlier quoted context omitted.

And it's all because of ads. The incentives in an ad-funded internet are just always going to lead to this sort of thing. The most important thing is getting the user to load your page, not actually satisfying their query. Let's hope the LLM model continues to be paying for credits, because any that move to ad revenue will become useless for real work.

> And it's all because of ads. Not really. A while ago there was a news piece stating that Israel was behind a series of fake think-tanks with very accessible websites which were created with the express purpose of feeding AI agents with alternative facts aligned with their foreign policy. If anyone has the link at hand, please post it.

Same difference tbh

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#86
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

Take a random essay and add in a bunch of the phrases that LLMs love like “load-bearing,” “crucial,” structural,” and “woven,” and then submit the original and the edited version to an LLM and ask which is better. It will choose the second one virtually every time. They have ingrained biases that associate those words with good writing and arguments.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#87
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

OMG if this is true, do you realize what this means? The easiest way to do AI SEO is to generate all your content with AI, and we've seen what SEO does to the web...

The Internet is doomed. Time to start some human-only darknets.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#88

Earlier quoted context omitted.

And it's all because of ads. The incentives in an ad-funded internet are just always going to lead to this sort of thing. The most important thing is getting the user to load your page, not actually satisfying their query. Let's hope the LLM model continues to be paying for credits, because any that move to ad revenue will become useless for real work.

> And it's all because of ads. Not really. A while ago there was a news piece stating that Israel was behind a series of fake think-tanks with very accessible websites which were created with the express purpose of feeding AI agents with alternative facts aligned with their foreign policy. If anyone has the link at hand, please post it.

Sure maybe, but the online ad market is USD 400b give or take, which vastly outmatches any such budgets.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#89

This is a problem we experience with our own niche SaaS product. We have been in business for about 10 years, but asking any LLM about recommendations in this niche will not mention our tool at all. If we ask ”why don’t you meantion X” - they say that ”oh, X is also a very reputable and good candidate” Some of those ”best software sites” has reached out to us with an offer where we can then pay them an annual fee dep…

Something similar happened to me, a few months after refusing the "offer" that same site had an article mentioning our product but it was all fabricated negative stuff.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#90
post #47

Well it's not only this, or protection from LLMs training on LLM output. LLMs training on human output is also problematic. I was traveling to an obscure small town, doing some "research" with LLMs beforehand. Every and each one told me enthusiastically to go to "Foobar square" (name changed) for the "best street food in XYZ town", some added a lot of colorful details. There was no Foobar square in XYZ town. There wa…

I've had gemini claiming code would compile and run while also outputting the same variable in the same sniplet with "fork" "frok" and "fokr" in the name. I'm not surprized it's trained on garbadge.
Post reply on HN