Live data from Hacker News

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

trellner.com

201–210 of 265 posts

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#201

I do think models currently don't have enough source skepticism. If you look at agent traces when asked to compare two options to help inform a decision, many of the comparison pages cited in research are often hosted by one of the companies being compared; nearly all are AI-generated AEO plays. Not deeply considering the motive of published information is currently a glitch that can be exploited, but the window will…

There are a lot of ways to make this a lot better easily. First of all, they could use a blacklist of sites that sell guest posts on adsy/etc.Also, if an article only links to one of the products listed, or only one is a dofollow link, they should also be excluded.

That'd probably cut down on a huge portion of spam by itself.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#202

Earlier quoted context omitted.

SEO is what ruined the web AFAIC. > Time to start some human-only darknets. I know very little about darknets. How could you ensure that they are human-only?

Lose anonymity and bring back key-signing parties. Maybe you can't guarantee that everything is human-generated, but at least you know the chain of trust that leads to the human that signed off. Yes, I'm aware of the irony of creating a darknet that only works by removing anonymity.

I think an invite system like lobste.rs could be enough. Prune bots at common ancestors

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#203

I've been vary of using ai to search considering all the spam out there. I think I'd rather, perhaps naively, whitelist wikipedia, reddit, arxiv, some news sources, etc than include everything. Is there nothing out there that does this? I'm paying for kagi and I can see that it has an api, is that maybe sufficient if configured properly?

Reddit is full of ai accounts now tho

Yep. Anything, but in particular less popular subreddits, are absolutely infested. Sometimes 80% of the responses I get are bots.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#204
post #171
post #18

If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated we…

I noticed this when I tried to get LLMs to play text adventures. Early on in my experiments, I wanted to give them hints when they got stuck at a puzzle. I did this by stopping the loop, injecting thoughts into the LLMs own persistent scratchpad (as if it had thought of that itself), and then starting the loop up again.[1] That way, the LLM would read what it had intended to remember from the previous turn including…

Imagine telling a person something that goes contrary to their conditioning, to their beliefs. They would tend to dismiss it completely and choose not to take it into account, even while knowing it's true.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#205

Honestly, I will just flag every post that is entirely AI slop from now on. This has to stop. The home page for this "independent research firm" is also 100% nonsense [1]. "The record a machine reads is not the one a company writes.". Ironically this low-effort spam is exactly what this report warns about, and does not belong in HN - or anywhere else. [1] https://trellner.com/

Unfortunately a large amount of people on this particular site love this, and they will meet your disgust with equal enthusiasm. Many people on here fancy themselves kindred spirits with the most ghoulish VCs you could imagine, and to them, an AI filled internet is a sign the system is working, and approval is a chance to be part of the elite who "get it."

Nuance is challenging for the general public when it comes to AI adoption.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#206

I do think models currently don't have enough source skepticism. If you look at agent traces when asked to compare two options to help inform a decision, many of the comparison pages cited in research are often hosted by one of the companies being compared; nearly all are AI-generated AEO plays. Not deeply considering the motive of published information is currently a glitch that can be exploited, but the window will…

Sounds like a tricky problem that will get a low-tech solution like a blacklist, whitelist, or chatGPT-approved vendor list.

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#207

Earlier quoted context omitted.

Your one example doesn't make all of LLMs a lie. It's like someone reading a National Enquirer article about "Hillary Clinton being an alien from outer space" (a real headline topic from decades ago) and drawling the conclusion that all journalism is "a lie". The user has to understand media literacy and be at least a little skeptical of the claims that are made, then cross reference with another source.

It kind of does, mathematically. It doesn't label confidence. If 0.0001% of answers is a lie, without knowing which parts are a lie exactly, you cannot trust any of them. If you need to independently verify every fact, why not just gather facts yourself in the first place. Let's say, a mathematical concept of lie. I still use them every day, of course.

"why not just gather facts yourself" vs "i use them every day", the duality of man. But yes, I appreciate that LLM output is theoretically completely untrustworthy- but when in practice I observe that it's around 90% accurate, I have to rely on my own internal calibration for how useful it is (depends on type, nature of task ofc)

Re: Three sites made 215,128 “best software” pages for AI. Perplexity cites them

#209

Earlier quoted context omitted.

It kind of does, mathematically. It doesn't label confidence. If 0.0001% of answers is a lie, without knowing which parts are a lie exactly, you cannot trust any of them. If you need to independently verify every fact, why not just gather facts yourself in the first place. Let's say, a mathematical concept of lie. I still use them every day, of course.

"why not just gather facts yourself" vs "i use them every day", the duality of man. But yes, I appreciate that LLM output is theoretically completely untrustworthy- but when in practice I observe that it's around 90% accurate, I have to rely on my own internal calibration for how useful it is (depends on type, nature of task ofc)

As they say, the less you know, the better LLM output is. :D
Post reply on HN