Live data from Hacker News

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

trellner.com

51–60 of 265 posts

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#51
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

>It always picks its own

Makes sense to me, in that its own output would align closer to its own training set

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#52

Honestly, I will just flag every post that is entirely AI slop from now on. This has to stop. The home page for this "independent research firm" is also 100% nonsense [1]. "The record a machine reads is not the one a company writes.". Ironically this low-effort spam is exactly what this report warns about, and does not belong in HN - or anywhere else. [1] https://trellner.com/

Also, that user's last four (three of them in the last hour) submissions have all been similar "finding" reports from a Claude-generated mystery research group website. All which contain exclusively AI slop articles. Ugh.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#54
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

The other day, I remember an article was posted to HN about something, but it came from a company that provides SEO services to companies by doing something like this: 1. For a given company, analyze their target audiences and the questions they are likely to ask LLMs about. 2. For each such question, ask it to each of the major LLMs, and compute the KL divergence between the pages they want to rank for the question…

And it's all because of ads. The incentives in an ad-funded internet are just always going to lead to this sort of thing. The most important thing is getting the user to load your page, not actually satisfying their query.

Let's hope the LLM model continues to be paying for credits, because any that move to ad revenue will become useless for real work.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#55

Earlier quoted context omitted.

Well, was. They did a bad job of it these last few years which allowed any AI that could crawl the web to seem amazing for search because it could pick the best posts from Reddit or whatever other forum had the best context for your question, but now we're watching the AI snake eat its own tail.

Hard to conclusively beat spam when your primary means of making money is selling the very ads that the spammers are using to make money.

Theoretically, Goggle doesn't care which websites run their ads, so they might as well give you the most useful ones. Search doesn't really work for engagementmaxxing.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#57

Earlier quoted context omitted.

Hard to conclusively beat spam when your primary means of making money is selling the very ads that the spammers are using to make money.

Theoretically, Goggle doesn't care which websites run their ads, so they might as well give you the most useful ones. Search doesn't really work for engagementmaxxing.

If you send people to the optimal website containing exactly the information they are after, then you get fewer ad impressions than if you send them to a suboptimal website that has them going back and clicking on more links.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#58
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

The other day, I remember an article was posted to HN about something, but it came from a company that provides SEO services to companies by doing something like this: 1. For a given company, analyze their target audiences and the questions they are likely to ask LLMs about. 2. For each such question, ask it to each of the major LLMs, and compute the KL divergence between the pages they want to rank for the question…

There is a term for the general practice of optimizing responses called GEO - Generative Engine Optimization [0]. A cousin of SEO and equally unsavory in how trust is being eroded through info shaping. Self-discovery by individuals is the victim.

0: https://en.wikipedia.org/wiki/Generative_engine_optimization

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#59
post #47

Well it's not only this, or protection from LLMs training on LLM output. LLMs training on human output is also problematic. I was traveling to an obscure small town, doing some "research" with LLMs beforehand. Every and each one told me enthusiastically to go to "Foobar square" (name changed) for the "best street food in XYZ town", some added a lot of colorful details. There was no Foobar square in XYZ town. There wa…

I think this was a common game on city/town subs. It happened here, there was a post asking for a good restaurant and someone just made up a name. It went viral and people started posting made-up menus for the place, reviews, and for a couple of months any time someone asked about a restaurant this fictional place would get mentioned.

It was all done as a joke to see if they could get Gemini or ChatGPT to start recommending it.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#60
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

Well, the one it generated is based on how it thought the best way to solve the problem was.

I am sure most humans would pick code written in their style, too.

Post reply on HN