Live data from Hacker News

SlopStop: Community-driven AI slop detection in Kagi Search

blog.kagi.com

161–170 of 271 posts

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#161

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

That's basically how "AI detectors" work, they're just ML models trained to classify human- vs LLM-generated content apart. As we all (hopefully) know, despite provider claims, they don't really work any well.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#162
post #88

So we have two universes. One is pushing generated content up our throats - from social media to operating systems - and another universe where people actively decide not to have anything to do with it. I wonder where the obstinacy on the part of certain CEOs come from. It's clear that although such content does have its fans (mostly grouped in communities), people at large just hate arificially-generated content. We…

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

Remains to be seen if that's sustainable or a flash in the pan.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#163

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

That's basically the entire idea behind GANs - Generative AI.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#164

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

There's already methods that attempt that.

It works for images because diffusion models leave artifacts, but doesn't work so well for text.

Text is an incredibly information dense data format. The diffusion artifacts kind of sneaks into the "extra data" in an image.

The other part is that GPT style models are effectively explicitly trained to minimize that entropy you're mentioning.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#165

I wish a smarter person would research or comment on this theory I have: Training a model to measure the entropy of human generated content vs LLM generated content might be the best approach to detecting LLM generated content. Consider the "will smith eating spaghetti test", if you compare the entropy (not similarity) between that and will smith actually eating spaghetti, I naively expect the main difference would b…

That's basically how "AI detectors" work, they're just ML models trained to classify human- vs LLM-generated content apart. As we all (hopefully) know, despite provider claims, they don't really work any well.

Correct, hence slopstop leveraging other signals than just the content

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#166

Definitely anecdata but an eye opener for me: I've been using Anthropic's models with gptel on Emacs for the past few months. It has been amazing for overviews and literature review on topics I am less familiar with. Surprisingly (for me) just slightly playing with system prompts immediately creates a writing style and voice that matches what _I_ would expect from a flesh agent. We're naturally biased to believe our…

That's definitely true, but keep in mind the economics of cranking out AI slop. The whole point is that you tell it "yo ChatGPT, write 1,000 articles about knitting / gardening / electronics and organize them into a website". You then upload it to a server and spend the rest of the day rolling in $100 bills. If you spend days or weeks fine-tuning prompts to strike the right tone, reviewing the output for accuracy, et…

https://xkcd.com/810/

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#167
post #88

So we have two universes. One is pushing generated content up our throats - from social media to operating systems - and another universe where people actively decide not to have anything to do with it. I wonder where the obstinacy on the part of certain CEOs come from. It's clear that although such content does have its fans (mostly grouped in communities), people at large just hate arificially-generated content. We…

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

Dude I had to stop watching that “sleepy whatever” channel. It was so blatant simply based upon how frequent the “thing” was posting. It’s simply not possible for a human to crank out well researched two hour long videos daily. And even then, the things content is so repetitive in each video (granted that might be the point, it is “sleepless historian” after all).

That sleepless channel is one of an entire series of very similar channels with the same voice and same “style” of content. Some get lots of views, others not so much.

Honestly, eventually people will spot that shit stuff from a mile away. None of it is unique nor does it add any “entropy” as some other commenter here said.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#168
post #152

Earlier quoted context omitted.

I think search engines should be worried, because people will silently lose faith in their results and start using AI chat instead. If search engines fail to find genuine, authentic content for me, and they just pipe me to LLM articles, I may as as well go straight to the LLM.

That will, if it is really adopted that widely, result in a freeze on available information.

Why would anybody even bother publishing or adding new content if the only thing that ever reads or interacts with it are bots?

I use the shit out LLM’s but you know what they can’t do? Create brand new ideas. They can refine yours, sure. They can take existing knowledge and map it into whatever you’re cooking. But on their own, nope. They just repeat what is in their training data and context window.

If all “new” content comes from LLM’s drawing from a huge pool of other LLM content… it’s just one giant echo chamber with nothing new being added. A planet wide circle jerk of LLMs complementing each other on what excellent ideas they all have and how they are really cutting to the heart of the issue. “Now I see the issue” they all say based on the slop context ingested from some other LLM who “saw the issue” from a third LLM. It’s LLMs all the way down.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#169
post #154
post #38

Earlier quoted context omitted.

Should we care? It's a tool. If you can manage to make it look original, then what can we do about it? Eventually you won't be able to detect it.

Objectively we should care because the content is not the whole value proposition of a blog post. The authenticity and trust of validity of the content comes from your connection to the human that made it. I don't need to fact check a ride review from an author I trust, if they actually ride mountain bikes. An AI article about mountain bikes lacks that implicit trust and authenticity. The AI has never ridden a bike b…

“It's worked on all the MSI boards I have tried it on.”

I love when they do that. It’s like a glitch in the matrix. It snaps you out of the illusion that these things are more than just a highly compressed form of internet text.

Re: SlopStop: Community-driven AI slop detection in Kagi Search

#170
post #88

So we have two universes. One is pushing generated content up our throats - from social media to operating systems - and another universe where people actively decide not to have anything to do with it. I wonder where the obstinacy on the part of certain CEOs come from. It's clear that although such content does have its fans (mostly grouped in communities), people at large just hate arificially-generated content. We…

There is a huge audience for AI-generated content on YouTube, though admittedly many of them are oblivious to the fact that they are watching AI-generated content. Here are several examples of videos with 1 million views that people don't seem to realize are AI-generated: * https://www.youtube.com/watch?v=vxvTjrsNtxA * https://www.youtube.com/watch?v=KfDnMpuSYic These videos do have some editing which I believe was d…

I really, really wish that Youtube would start tagging this category of video to increase the visibility to end users. My feeling is that the main reason this content might be "winning" in the market is the sheer volume.
Post reply on HN