Live data from Hacker News

SlopStop: Community-driven AI slop detection in Kagi Search

blog.kagi.com

141–150 of 271 posts

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#141

Isn't the scalable approach to ask AI to identify AI (and have a human review the results, but that's required no matter what)? I also doubt most people will be able to detect AI text generated with a non-default "voice" in the prompt.

AI is unreliable at detecting AI or else this would be a trivial problem to solve.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#142

Earlier quoted context omitted.

I just requested access to the database @freediver so hopefully it should be integrated into https://hcker.news soon. I appreciate Kagi's community-driven approach. The open Small Web list[0] is invaluable. Applying a smallweb filter[1] on HN brings a breath of fresh air to the frontpage. 0: https://github.com/kagisearch/smallweb 1: https://hcker.news/?smallweb=true

I like the effort, but it's super restrictive. They exclude all of Substack on principle (but weirdly, allow blogspot.com and wordpress.com). They exclude anything that isn't a blog. And they exclude blogs that aren't updated often enough. The end result is that there's a lot of "small web" stuff that doesn't show up. Looking at my bookmarks, I think 90% of them are in the "small web" category in spirit, but maybe 10…

I understand the substack exclusion. The paywall is not user friendly.

If you don't mind, it'd be cool to take a look at your bookmark domains so that I could potentially augment the filter on my site. If you're interested, my email is in bio.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#143

So we have two universes. One is pushing generated content up our throats - from social media to operating systems - and another universe where people actively decide not to have anything to do with it. I wonder where the obstinacy on the part of certain CEOs come from. It's clear that although such content does have its fans (mostly grouped in communities), people at large just hate arificially-generated content. We…

If creators are required to disclose that they used AI to create, modify, or manipulate content then I should be able to filter out content created with AI. Even if I'm thinking of a specific video it's getting harder to find things because of the ridiculous amount of mass-produced slop out there.

I don't really care if people produce this sort of crap; let the market sort it out, maybe something of value will come of it. It's the fact that, as Kagi points out, it's getting more and more difficult to produce anything of value because content creators operating in good faith with good intentions get drowned out by slop peddlers who have no such limitations or morals.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#144
I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content.

Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would be entropy. when we say something looks "real" I think we're just talking about our expectation of entropy for that scene. An LLM can detect that it is a person eating a spaghetti see what the entropy is compared to the entropy it expects for the scene based on its training. In other words, train a model with specific entropy measurements along side actual training data.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#145
post #38

Earlier quoted context omitted.

> Even now if you put an effort into prompting and context building, you can achieve 100% human like results. Are we personally comfortable with such an approach? For example, if you discover your favorite blogger doing this.

Should we care? It's a tool. If you can manage to make it look original, then what can we do about it? Eventually you won't be able to detect it.

We should care if it is lower in quality than something made by humans (e.g. less accurate, less insightful, less creative, etc.) but looks like human content. In that scenario, AI slop could easily flood out meaningful content.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#146
post #60
post #57

Earlier quoted context omitted.

Explicitly in the article, one of the headings is "AI slop is deceptive or low-value AI-generated content, created to manipulate ranking or attention rather than help the reader." So yes, they are proposing marking bad AI content (from the user's perspective), not all AI-generated content.

Which troubles me a bit, as 'bad' does not have same definition for everyone.

A simple definition would be: Its bad if it isn't labeled as AI content or if there is not a mechanism that allows you to filter out AI content.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#148

Earlier quoted context omitted.

Slop is about thoughtless use of a model to generate output. Output from your paper's model would still qualify as slop in our book. Even if your model scored extremely high perplexity on an LLM evaluation we'd likely still tag it as slop because most of our text slop detection is using sidechannel signals to parse out how it was used rather than just using an LLM's statistical properties on the text.

Would love to see proof of this claim that you can tag antislopped LLM text as LLM generated. I'm willing to bet money that you can't.

If its not labeled as generated by AI, then that in of itself makes it deceptive and therefore slop.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#149

Earlier quoted context omitted.

Hot take but I don't care if the content I consume is AI-generated or not. First of all, while sometimes I need high-effort quality content, sometimes I want my brain to rest and then AI-generated slop is completely okay. He who didn't binge-watch garbage reality TV can cast the first stone. Second, just because something is AI-generated it doesn't automatically mean it's slop, just like human-generated content isn't…

> He who didn't binge-watch garbage reality TV can cast the first stone Stand by then, because I have rocks and according to you, licence to throw them. You are free to watch all the slop you want. All I want is for your slop, to not be at the cost of all other media and content. Have a SlopTube, have SlopFlix, go for it! But do it in a way that is _separate_ and doesn’t inflict it on the rest of us, who would _like_…

Just let me choose a filter when I'm doing a search on YouTube and that's a good start. Beyond that I can just block or 'don't recommend this channel' for anything that shows up in my feed, but the fact that these platforms don't let people say 'I don't want this garbage' is the biggest issue I have with it.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#150

Earlier quoted context omitted.

Hot take but I don't care if the content I consume is AI-generated or not. First of all, while sometimes I need high-effort quality content, sometimes I want my brain to rest and then AI-generated slop is completely okay. He who didn't binge-watch garbage reality TV can cast the first stone. Second, just because something is AI-generated it doesn't automatically mean it's slop, just like human-generated content isn't…

> He who didn't binge-watch garbage reality TV can cast the first stone Stand by then, because I have rocks and according to you, licence to throw them. You are free to watch all the slop you want. All I want is for your slop, to not be at the cost of all other media and content. Have a SlopTube, have SlopFlix, go for it! But do it in a way that is _separate_ and doesn’t inflict it on the rest of us, who would _like_…

Your later point is hard to convey to people who don't want to hear it.

I don't want AI content, even if it is as good, or even if it were better. The human element IS the point, not an implementation detail.

An AI song about sailing at sea is meaningless because I know the AI has never sailed at sea. This is a standard we hold humans to, authenticity is important even for human artists, why would we give AI a pass on it?

And I mean this earnestly, if an AI in a corporeal form really did go sailing, I might then be interested in its song about sailing.

Post reply on HN