Live data from Hacker News

SlopStop: Community-driven AI slop detection in Kagi Search

blog.kagi.com

131–140 of 271 posts

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#133

Earlier quoted context omitted.

Slop is about thoughtless use of a model to generate output. Output from your paper's model would still qualify as slop in our book. Even if your model scored extremely high perplexity on an LLM evaluation we'd likely still tag it as slop because most of our text slop detection is using sidechannel signals to parse out how it was used rather than just using an LLM's statistical properties on the text.

Would love to see proof of this claim that you can tag antislopped LLM text as LLM generated. I'm willing to bet money that you can't.

Here's what pattern suppression actually does on a model that's trained to open its writing with "You're absolutely right.":

You're spot-on. You're bang-on. You're dead right. You're 100% correct. I couldn't agree more. I agree completely. That's exactly right. That's absolutely correct. That's on the nose. You hit the nail on the head. Right you are. Very true. Exactly — well said. Precisely so. No argument from me. I'll second that. I'm with you 100%. You've got it exactly. You've hit the mark. Affirmative — that's right. Unquestionably correct. Without a doubt, you're right.

I'm willing to bet money you can easily tag these openers yourself.

This sampling strategy and the elaborate scheme to bake its behavior into the model during the post-training are terribly misguided, because they don't fix the underlying mode collapse. It's formulated as narrowing down the output distribution, but as with many things in LLMs it manifests itself on a much higher semantical level - during the RL (at least using the current methods) the model narrows the many-to-many mapping of high-level ideas that the pretrained model has down to one-to-one or even many-to-one. If you naively suppress repetitive n-grams that are not semantically aware and manually constructed patterns that don't scale, it will just slip out at the first chance, spamming you with minor non-repetitive variations of the same high-level idea.

You'll never have the actual semantic variety unless you fix mode collapse. Referencing n-grams or manually constructed regexes as a source of semantical diversity automatically makes the method invalid, no matter how elaborate your proxy is. I can't believe that after all this time you persist in this and don't see the obvious issue that's been pointed at multiple times.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#134

This is so, so exciting. I hope HN takes inspiration and adds a similar flag. :)

I just requested access to the database @freediver so hopefully it should be integrated into https://hcker.news soon.

I appreciate Kagi's community-driven approach. The open Small Web list[0] is invaluable. Applying a smallweb filter[1] on HN brings a breath of fresh air to the frontpage.

0: https://github.com/kagisearch/smallweb

1: https://hcker.news/?smallweb=true

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#135

We wrote the paper on how to deslop your language model: https://arxiv.org/abs/2510.15061

People don't call it slop because of repetitive patterns they call it slop because it's low-effort, uninsightful, meaningless content cranked out in large volumes

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#137

This is so, so exciting. I hope HN takes inspiration and adds a similar flag. :)

I just requested access to the database @freediver so hopefully it should be integrated into https://hcker.news soon. I appreciate Kagi's community-driven approach. The open Small Web list[0] is invaluable. Applying a smallweb filter[1] on HN brings a breath of fresh air to the frontpage. 0: https://github.com/kagisearch/smallweb 1: https://hcker.news/?smallweb=true

I like the effort, but it's super restrictive. They exclude all of Substack on principle (but weirdly, allow blogspot.com and wordpress.com). They exclude anything that isn't a blog. And they exclude blogs that aren't updated often enough.

The end result is that there's a lot of "small web" stuff that doesn't show up. Looking at my bookmarks, I think 90% of them are in the "small web" category in spirit, but maybe 10% have any chance of appearing on the Kagi list.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#138

Isn't the scalable approach to ask AI to identify AI (and have a human review the results, but that's required no matter what)? I also doubt most people will be able to detect AI text generated with a non-default "voice" in the prompt.

Asking AI to identify AI is like claiming that we will solve alignment by building "good" AI that beats "bad" AI.

Maybe it could work, but that seems like a chain of assumptions and hope that isn't particularly realistic.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#139

Definitely anecdata but an eye opener for me: I've been using Anthropic's models with gptel on Emacs for the past few months. It has been amazing for overviews and literature review on topics I am less familiar with. Surprisingly (for me) just slightly playing with system prompts immediately creates a writing style and voice that matches what _I_ would expect from a flesh agent. We're naturally biased to believe our…

That's definitely true, but keep in mind the economics of cranking out AI slop. The whole point is that you tell it "yo ChatGPT, write 1,000 articles about knitting / gardening / electronics and organize them into a website". You then upload it to a server and spend the rest of the day rolling in $100 bills.

If you spend days or weeks fine-tuning prompts to strike the right tone, reviewing the output for accuracy, etc, then pretty much by definition, you're undermining the economic benefits of slopification. And you might accidentally end up producing content that's actually insightful and useful, in which case, you know... maybe that's fine.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#140
post #88

Earlier quoted context omitted.

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

Hot take but I don't care if the content I consume is AI-generated or not. First of all, while sometimes I need high-effort quality content, sometimes I want my brain to rest and then AI-generated slop is completely okay. He who didn't binge-watch garbage reality TV can cast the first stone. Second, just because something is AI-generated it doesn't automatically mean it's slop, just like human-generated content isn't…

> He who didn't binge-watch garbage reality TV can cast the first stone

Stand by then, because I have rocks and according to you, licence to throw them.

You are free to watch all the slop you want. All I want is for your slop, to not be at the cost of all other media and content. Have a SlopTube, have SlopFlix, go for it! But do it in a way that is _separate_ and doesn’t inflict it on the rest of us, who would _like_ human produced content, even if the AI stuff is “just as good”.

Post reply on HN