Live data from Hacker News

12k AI-generated blog posts added in a single commit

github.com

141–150 of 167 posts

Re: 12k AI-generated blog posts added in a single commit

#141

I suspect we'll address this by just going back to older ranking algorithms for search. We'll go back to the primary signal of good content being links from trusted sources. People gaming the content based algorithms will eventually cause their own downfall.

This has been the status quo for more than a decade.

In the past SEO blogspam was done by cheap freelancers, and there were several agencies selling the service.

Experts identify blogspam quite easily, but laypeople eat it up and use as reference in conversations and to make decisions.

Google has known about it, has been in contact with such agencies and companies, and has been refusing to do anything about it for the longest time.

Re: 12k AI-generated blog posts added in a single commit

#142

It's becoming much harder to determine on a daily basis what content is original, thought-out by a person, and trustworthy. Ironically, verifiably-old content is easier to trust now. Examples from recent personal experience: 1) Some time ago I was searching for growing information about a specific and uncommonly-grown plant, and was led to a top-ranked website with long pages containing everything about it, including…

You're making the classic mistake of looking for a trustworthy information source and then trusting it, instead of focusing on whether the information itself is trustworthy regardless of source. It's literally the same as my grandma saying "they said so on TV, therefore it must be true" while completely dismissing anything I've read on the internet because reasons. If you develop the skill of judging information by i…

Information itself cannot be trustworthy. It can be right, it can be wrong, or it can be somewhere in between. Only a source can have trustworthiness, as it's a mixed measure of reputation and provable accuracy.

You filter out known untrustworthy sources to not waste your time verifying false information 100x more than you need to. I know The Onion is a satire publication. I do not need to verify its claims. It's an intentionally untrustworthy source. I know that LLMs can hallucinate information, so I verify with a more trustworthy source. I cross-reference things random people say on the internet, because random people on the internet are not, individually, trustworthy sources of information.

If a rocket engineer explains to me why Rocket A isn't flight ready, I'm more inclined to believe them than if a random commenter on the internet explains it to me. Because the one source is more trustworthy than another, and if I wanted to verify the claim myself I'd have to spend a lot of time studying rocket science.

Re: 12k AI-generated blog posts added in a single commit

#143

Earlier quoted context omitted.

You're making the classic mistake of looking for a trustworthy information source and then trusting it, instead of focusing on whether the information itself is trustworthy regardless of source. It's literally the same as my grandma saying "they said so on TV, therefore it must be true" while completely dismissing anything I've read on the internet because reasons. If you develop the skill of judging information by i…

No it's not the same as your grandma. The point is that it's now more expensive to find the correct information to learn from. You don't know it's an LLM ahead of time, and you may spend hours until you figure out something is off. Hence why reputable sources will become more valuable. > If you develop the skill of judging information by its merit rather than source.. Did you read example #1? I'm not talking about so…

What if your physics book is wrong because knowledge has advanced since it was released - you can still find lots of publications and people with degrees blissfully unaware of Hawking Radiation. What if your botanical book is wrong because facts have changed since then - climate is changing and so does flora. What if your book is wrong because it's state-funded propaganda mixed with petty fights of a bunch of people with suits and strong opinions disguised as academia - a huge chunk of linguistics is dealing with exactly this issue.

Again, you seem to miss the point that the idea of questioning new information, which was already useful to navigate life before LLMs, before television, before newspaper, before print, before clay tablets, even before speech itself, is equally applicable to LLMs as to any other form of communication. You just need to upgrade your strategies a little and that's it. Don't blow this out of proportion "somebody gasp lied to me on the internet!".

Re: 12k AI-generated blog posts added in a single commit

#145

It's becoming much harder to determine on a daily basis what content is original, thought-out by a person, and trustworthy. Ironically, verifiably-old content is easier to trust now. Examples from recent personal experience: 1) Some time ago I was searching for growing information about a specific and uncommonly-grown plant, and was led to a top-ranked website with long pages containing everything about it, including…

You're making the classic mistake of looking for a trustworthy information source and then trusting it, instead of focusing on whether the information itself is trustworthy regardless of source. It's literally the same as my grandma saying "they said so on TV, therefore it must be true" while completely dismissing anything I've read on the internet because reasons. If you develop the skill of judging information by i…

You do ultimately need to trust some sources to some degree. You can try to cross-correlate multiple sources (and this is in general a good habit!), but that depends on some level of trustworthiness in the sources you are looking at, you're not at all immune to misinformation by doing this (especially if multiple sources are, undisclosed, being generated from the same LLM. You can also get citenogenesis even pre-LLMs). And of course for some things it's possible to try to verify directly yourself, but this is infeasible to do for everything you depend on.

Re: 12k AI-generated blog posts added in a single commit

#149
post #64

Earlier quoted context omitted.

You should care because this website has a high ranking on Google and these 12000 posts will show up every time you search something programming related.

Stop using Google. Kagi lets you block and prioritize sites.

>Stop using Google

I've been using DuckDuckGo for years now, but their search results have now become so terrible, it's nearly unusable. And I have to "quote" every word, otherwise it just randomly omits it from search for no reason. Honestly it's so bad I wonder if all the developers left, and the site is just coasting along.

Maybe I'll try Kagi, but it's not something I can pitch to normies.

Re: 12k AI-generated blog posts added in a single commit

#150
post #17

I am so glad DuckDuckGo allows blocking specific sites from the search. Just did this for a domain linked in this repository.

It would be nice if DDG made even a token attempt at making their search not shit. I still use it, but mostly out of habit and because I suspect every alternative is also shit.
Post reply on HN