Earlier quoted context omitted.
Take a random essay and add in a bunch of the phrases that LLMs love like “load-bearing,” “crucial,” structural,” and “woven,” and then submit the original and the edited version to an LLM and ask which is better. It will choose the second one virtually every time. They have ingrained biases that associate those words with good writing and arguments.
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
141–150 of 265 posts
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#142Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#143Well it's not only this, or protection from LLMs training on LLM output. LLMs training on human output is also problematic. I was traveling to an obscure small town, doing some "research" with LLMs beforehand. Every and each one told me enthusiastically to go to "Foobar square" (name changed) for the "best street food in XYZ town", some added a lot of colorful details. There was no Foobar square in XYZ town. There wa…
Your one example doesn't make all of LLMs a lie. It's like someone reading a National Enquirer article about "Hillary Clinton being an alien from outer space" (a real headline topic from decades ago) and drawling the conclusion that all journalism is "a lie". The user has to understand media literacy and be at least a little skeptical of the claims that are made, then cross reference with another source.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#144Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#145Earlier quoted context omitted.
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#146Earlier quoted context omitted.
> And it's all because of ads. Not really. A while ago there was a news piece stating that Israel was behind a series of fake think-tanks with very accessible websites which were created with the express purpose of feeding AI agents with alternative facts aligned with their foreign policy. If anyone has the link at hand, please post it.
Sure maybe, but the online ad market is USD 400b give or take, which vastly outmatches any such budgets.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#147Earlier quoted context omitted.
Your one example doesn't make all of LLMs a lie. It's like someone reading a National Enquirer article about "Hillary Clinton being an alien from outer space" (a real headline topic from decades ago) and drawling the conclusion that all journalism is "a lie". The user has to understand media literacy and be at least a little skeptical of the claims that are made, then cross reference with another source.
Okay, but the anecdote states that every model repeated the pseudo-factoid about Foobar square, not just the 4 GB open source model equivalent of a tabloid.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#148Earlier quoted context omitted.
Okay, but the anecdote states that every model repeated the pseudo-factoid about Foobar square, not just the 4 GB open source model equivalent of a tabloid.
Without disclosing what you were prompting for, it's impossible to evaluate your claim.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#149If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…
That makes sense. What an LLM does is output what the model thinks is the best set of tokens in response to a given input, so when you ask it to judge the best response to that input it is going to conclude that the best one is the one that must closely matches what it would output, which is what it did output.
Of course you aren't giving exactly the same context+input, but close enough that any difference doesn't push the output it made far from what it is going to say is ideal.
Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them
#150Begun, the AI SEO wars have.
It does seem rather impactful. I've seen an extremely aggressive uptick in API key requests and sales that I'm not sure where it's coming from. Like it's up 5x over the summer. Been a bit confused about this since I do basically zero traditional marketing or SEO, but I think it's AI search tools that's suggesting my services.