Live data from Hacker News

SlopStop: Community-driven AI slop detection in Kagi Search

blog.kagi.com

211–220 of 271 posts

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#211

Earlier quoted context omitted.

Substack is definitely outside my idea of what “the small web” means (I realise this isn’t well defined and will mean different things to different people, though). It’s a platform and social network of sorts, rather than a neutral hosting provider and it’s too often used in a way that’s inauthentically commercial IMO.

Note that this is the admission policy for a per-blog whitelist - we're not talking about including *.substack.com as a "good" domain, just allowing someone to propose the inclusion of hacker-bob.substack.com. And the policy already allows wordpress.com or blogspot.com (the latter is probably mostly spam nowadays, with a few holdouts who have been using it for 20 years). Also note that Small Web allows YouTube channe…

> And the policy already allows wordpress.com or blogspot.com (the latter is probably mostly spam nowadays, with a few holdouts who have been using it for 20 years).

Do you mean the entire .wordpress.com and .blogspot.com are allowed as per the grandparent comment implies, or just individual blogs may or may no be allowed, exactly like substack?

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#212

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

That's basically how "AI detectors" work, they're just ML models trained to classify human- vs LLM-generated content apart. As we all (hopefully) know, despite provider claims, they don't really work any well.

In a non-adversial context (so when the author isn't disclosing it, but also not actively trying to hide it), AI image detection is giving me great results.

I think (currently) the problems are more about text, or post processing of other media to hide AI.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#213

What about human slop? start with HN a significant number of comments are pretty dire.

You must not use Kagi because a "human slop" system is available on both Kagi and HN. It's called a downvote and the article has an image how you can downvote links in search results. Just an FYI why you're getting downvoted for posting a dire comment yourself.

Then people can downvote AI slop or Human slop equally. Why do we need to discriminate against digital intelligence which is often leagues above the average mouth breathing Joe.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#214

Isn't the scalable approach to ask AI to identify AI (and have a human review the results, but that's required no matter what)? I also doubt most people will be able to detect AI text generated with a non-default "voice" in the prompt.

The next model will be trained away from samples that classify as AI and the cycle will go on. LLMs are good at things like that. People do that on purpose to match a given style or type of behaviour https://en.wikipedia.org/wiki/Generative_adversarial_network

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#215

Seems like they are equating all generated content with slop. Is that how people actually understand "slop"? https://help.kagi.com/kagi/features/slopstop.html#what-is-co... > We evaluate the channel; if the majority of its content is AI‑generated, the channel is flagged as AI slop and downranked. What about, y'know, good generated content like Neural Viz? https://www.youtube.com/@NeuralViz

> What about, y'know, good generated content like Neural Viz? There is no good AI generated content. I just clicked around randomly on a few of those videos and then there was this guy dual-wielding mice: https://youtu.be/1Ijs1Z2fWQQ?si=9X0y6AGyK_5Gaiko&t=19

> There is no good AI generated content.

What's good or bad is subjective. I've seen plenty of (in my opinion) good AI-generated content. But making such a sweeping statement suggests to me that your mind is made up on the topic.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#216

Earlier quoted context omitted.

I like the effort, but it's super restrictive. They exclude all of Substack on principle (but weirdly, allow blogspot.com and wordpress.com). They exclude anything that isn't a blog. And they exclude blogs that aren't updated often enough. The end result is that there's a lot of "small web" stuff that doesn't show up. Looking at my bookmarks, I think 90% of them are in the "small web" category in spirit, but maybe 10…

Substack is definitely outside my idea of what “the small web” means (I realise this isn’t well defined and will mean different things to different people, though). It’s a platform and social network of sorts, rather than a neutral hosting provider and it’s too often used in a way that’s inauthentically commercial IMO.

The social network seems relevant to me. It feels like people are posting for clout, trying to get to as many inboxes as possible, so they post a lot of marketing slop, just like on LinkedIn.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#217

Earlier quoted context omitted.

Thing is LLM can be trained on things not available on the public internet, unlike search engines having to return public URLs. I’m sure all these agentic AI are slurping in all proprietary codebases and their documentations and training on them. It’s the only way to one up the competition.

That is the part which most frightens me. If the LLMs become the bastions of truth because the open web has fallen to slop, truth becomes privatized and unverifiable.

There is a fair chance that we are already in the beginning of this. The web allowed AI to be jumpstarted because it made a lot of information available. Now the AI peddlers are incentivized to destroy the web so they have a monopoly.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#218

Earlier quoted context omitted.

That will, if it is really adopted that widely, result in a freeze on available information.

Why would anybody even bother publishing or adding new content if the only thing that ever reads or interacts with it are bots? I use the shit out LLM’s but you know what they can’t do? Create brand new ideas. They can refine yours, sure. They can take existing knowledge and map it into whatever you’re cooking. But on their own, nope. They just repeat what is in their training data and context window. If all “new” co…

Yes, it is very much a one-way street and the value of original creation is reduced to nil because they just steal it.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#219

Earlier quoted context omitted.

It is far worse. SEO spam was easy to detect for a human, even if it fooled the search engine. This is a proverbial deluge of crap and now you're left to find the crumbs. And the crap looks good. It's still crap, but it outperforms the real thing of look and feel as well as general language skills while it underperforms in the part that matters. But I can see why other search engines love it: it further allows them t…

> It's still crap, but it outperforms the real thing of look and feel as well as general language skills while it underperforms in the part that matters. The real thing meant human SEO spam? Or human writing?

Actual knowledge and understanding.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#220

Earlier quoted context omitted.

we just need human attestation. A vial of blood per comment

Isn't "Proof of Humanity" kind of interesting here: https://proofofhumanity.id

I'd want a "proof of humanity" without needing to reveal my identity...
Post reply on HN