Live data from Hacker News

SlopStop: Community-driven AI slop detection in Kagi Search

blog.kagi.com

191–200 of 271 posts

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#191

Earlier quoted context omitted.

I just requested access to the database @freediver so hopefully it should be integrated into https://hcker.news soon. I appreciate Kagi's community-driven approach. The open Small Web list[0] is invaluable. Applying a smallweb filter[1] on HN brings a breath of fresh air to the frontpage. 0: https://github.com/kagisearch/smallweb 1: https://hcker.news/?smallweb=true

I like the effort, but it's super restrictive. They exclude all of Substack on principle (but weirdly, allow blogspot.com and wordpress.com). They exclude anything that isn't a blog. And they exclude blogs that aren't updated often enough. The end result is that there's a lot of "small web" stuff that doesn't show up. Looking at my bookmarks, I think 90% of them are in the "small web" category in spirit, but maybe 10…

Substack is definitely outside my idea of what “the small web” means (I realise this isn’t well defined and will mean different things to different people, though).

It’s a platform and social network of sorts, rather than a neutral hosting provider and it’s too often used in a way that’s inauthentically commercial IMO.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#192

Earlier quoted context omitted.

That will, if it is really adopted that widely, result in a freeze on available information.

Thing is LLM can be trained on things not available on the public internet, unlike search engines having to return public URLs. I’m sure all these agentic AI are slurping in all proprietary codebases and their documentations and training on them. It’s the only way to one up the competition.

That is the part which most frightens me.

If the LLMs become the bastions of truth because the open web has fallen to slop, truth becomes privatized and unverifiable.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#193
>> Per our AI integration philosophy, we’re not against AI tools that enhance human creativity. But when it includes fake reviews, fabricated expertise, misinformation ...

There, the childish wish that you can control things the way you want to. Same as wishing that you can control which country gets the nukes. The wish that Tarzan is good and can be controlled to not to bring in humans, the wish that slaves help in work and can be controlled not to change demography, the wish that capitalism are good and can be controlled to avoid economic disparity and provide equality. When do we stop the children managing this planet?

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#194

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

That would flag poorly encoded videos too.

Another problem is AI generators will try to find “workaround”s to bypass this system. In theory sounds good, in practice I doubt it would work.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#195
post #88

So we have two universes. One is pushing generated content up our throats - from social media to operating systems - and another universe where people actively decide not to have anything to do with it. I wonder where the obstinacy on the part of certain CEOs come from. It's clear that although such content does have its fans (mostly grouped in communities), people at large just hate arificially-generated content. We…

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

I can attest to this.

I don't remember what channel but recently I have been into dexter and I have been watching a lot of dexter related content on youtube and I once think that I saw either down-right AI generated or very LLM-y style video / channel in general. Like, the way they speak etc. felt very AI generated imo.

Nobody questioned it in the comments.

I genuinely started wondering what is the point of AI generated content when people will find out its AI and then reject it or shame them etc. but I think that either I believed that humans in general would detect it more often or maybe the fact that people would start using AI in very sneaky ways maybe to not be labelled AI slop while still being very AI assisted.

I don't have problem with AI assistance but I just feel this hate when an AI generated voice speaks AI generated text which I recognize due to the patterns like

"It isn't just X, Its y" and the countless others examples.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#196
post #88

Earlier quoted context omitted.

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

Reddit has been full of bad fake stories for ages. All that AI does is automate it

Karma farming accounts I guess.

I loved it when sometimes on r/Aita or something people would call out the sheer inconsistencies of the karma farming accounts

"So you are telling me that you were 26 year old and now you are suddenly 40??"

Or just sheer inconsistencies which can make one laugh at the whole situation.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#197

"Begun, the slop wars have." I applaud any effort to stem the deluge of slop in search results. It's SEO spam all over again, but in a different package.

It is far worse. SEO spam was easy to detect for a human, even if it fooled the search engine. This is a proverbial deluge of crap and now you're left to find the crumbs. And the crap looks good. It's still crap, but it outperforms the real thing of look and feel as well as general language skills while it underperforms in the part that matters. But I can see why other search engines love it: it further allows them t…

> It's still crap, but it outperforms the real thing of look and feel as well as general language skills while it underperforms in the part that matters.

The real thing meant human SEO spam? Or human writing?

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#198

Earlier quoted context omitted.

I like the effort, but it's super restrictive. They exclude all of Substack on principle (but weirdly, allow blogspot.com and wordpress.com). They exclude anything that isn't a blog. And they exclude blogs that aren't updated often enough. The end result is that there's a lot of "small web" stuff that doesn't show up. Looking at my bookmarks, I think 90% of them are in the "small web" category in spirit, but maybe 10…

Substack is definitely outside my idea of what “the small web” means (I realise this isn’t well defined and will mean different things to different people, though). It’s a platform and social network of sorts, rather than a neutral hosting provider and it’s too often used in a way that’s inauthentically commercial IMO.

Note that this is the admission policy for a per-blog whitelist - we're not talking about including *.substack.com as a "good" domain, just allowing someone to propose the inclusion of hacker-bob.substack.com.

And the policy already allows wordpress.com or blogspot.com (the latter is probably mostly spam nowadays, with a few holdouts who have been using it for 20 years). Also note that Small Web allows YouTube channels under 400k subscribers (!). So it's really not that clean-cut.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#199

Earlier quoted context omitted.

Hot take but I don't care if the content I consume is AI-generated or not. First of all, while sometimes I need high-effort quality content, sometimes I want my brain to rest and then AI-generated slop is completely okay. He who didn't binge-watch garbage reality TV can cast the first stone. Second, just because something is AI-generated it doesn't automatically mean it's slop, just like human-generated content isn't…

> He who didn't binge-watch garbage reality TV can cast the first stone Stand by then, because I have rocks and according to you, licence to throw them. You are free to watch all the slop you want. All I want is for your slop, to not be at the cost of all other media and content. Have a SlopTube, have SlopFlix, go for it! But do it in a way that is _separate_ and doesn’t inflict it on the rest of us, who would _like_…

No, you get your separate HumanTube.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#200

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

I doubt AI slob is the solution of AI slob, far too error prone. Problem is we already had a slob advertising/attention economy, AI just made the problem more visible.

Any AI model can easily increase entropy by adding info bits and we would have a weird AI info war where people will become victims. If you consume info we deal with unknown spaghetti. Generating false info is too easy for a model.

Post reply on HN