Live data from Hacker News

When your hash becomes a string: Hunting Ruby's million-to-one memory bug

mensfeld.pl

61–70 of 70 posts

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#61
post #55

Earlier quoted context omitted.

I would never claim that we can reliably detect all AI generated text. There are many ways to write text with LLM assistance that is indistinguishable from human output. Moreover, models themselves are extremely bad at detecting AI-generated text, and it is relatively easy to edit these tells out if you know what to look for (one can try to prompt them out too, though success is more limited there). I am happy to mak…

On the other fork where I responded to your claims with a direct and detailed response, you insisted that my comment “isn't really that interesting” and just disengaged. I’m not going to write another detailed explanation of why your “slop === AI” premise is flawed. Go reread the other fork if you’ve decided you’re interested. > I find it interesting that you believe this claim is wildly conspirational I don’t believ…

Human experts can reliably detect some kinds of long-form, AI-generated text using exactly the same sorts of cues I've outlined: https://arxiv.org/html/2501.15654v1. You may take issue with the quality of the paper, but there have been very few studies like this and this one found an extremely strong effect.

I am making an even more limited claim than the article, which is only that it's possible for "experts" (i.e. people who frequently interact with LLMs as part of their day jobs) to identify AI generated text in long-form passages in a way that has very few false positives, not classify it perfectly. I've also introduced the caveat that this only applies to AI generated text that has received minimal or no prompting to "humanize" the writing style, not AI generated text in general.

If you would like to perform a higher-quality study with more recent models, feel free (it's only fair that I ask you to do an unreasonable amount of work here given that your argument appears to be that if I don't quit my lucrative programming job and go manually classify text for pennies on the dollar, it proves that it can't be done).

The reason this isn't offered as a service is because it makes no economic sense to do so using humans, not because it's impossible as you claim. This kind of "human" detection mechanism does not scale the way generation does. The cues that I rely on are also pretty easy to eliminate if you know someone is looking for them. This means that heuristics do not work reliably against someone actively trying to avoid human detection, or a human deliberately trying to sound like an LLM (I feel the need to reiterate this as many of the counterarguments to what I'm saying are to claims of this form).

> I’m not going to write another detailed explanation of why your “slop === AI” premise is flawed.

This isn't a claim that I made. I believe that text written with LLM assistance is not necessarily slop, and that slop is not necessarily AI generated. The only assertion I made regarding slop is that being written with LLM assistance with minimal prompting or editing is a strong predictor of slop, and that the heuristics I'm using (if present in large quantities) are a strong predictor of an article being written with LLM assistance with minimal prompting or editing. i.e. I, I am asserting that these kinds of heuristics work pretty well on articles generated by people who don't realize (or care) that there are LLM "tells" all over their work. The fact that many of the articles posted to HN are being accused of being LLM generated could certainly indicate that this is all just a massive witch hunt, but given the acknowledged popularity of ChatGPT among the general population and the fact that experts can pretty easily identify non-humanized articles, I think "a lot of people are using LLMs in the process of generating their blog posts, and some sizable fraction of those people didn't edit the output very much" is an equally compelling hypothesis.

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#62
post #55

Earlier quoted context omitted.

On the other fork where I responded to your claims with a direct and detailed response, you insisted that my comment “isn't really that interesting” and just disengaged. I’m not going to write another detailed explanation of why your “slop === AI” premise is flawed. Go reread the other fork if you’ve decided you’re interested. > I find it interesting that you believe this claim is wildly conspirational I don’t believ…

> believing that you (and whatever group you identify as “us”) are able to deterministically identify AI text I think you will find the OP said no such thing. They instead said they identified a mixture of writing styles consistent with a human author and an LLM. The OP says nothing about deterministically identifying LLMs, only that the style of specific sections is consistent with LLMs leading to the conclusion.

I think you find OP absolutely did say that.

> Parts of it were 100% LLM written. Like it or not, people can recognize LLM-generated text pretty easily

https://news.ycombinator.com/item?id=45868782

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#63
post #62

Earlier quoted context omitted.

> believing that you (and whatever group you identify as “us”) are able to deterministically identify AI text I think you will find the OP said no such thing. They instead said they identified a mixture of writing styles consistent with a human author and an LLM. The OP says nothing about deterministically identifying LLMs, only that the style of specific sections is consistent with LLMs leading to the conclusion.

I think you find OP absolutely did say that. > Parts of it were 100% LLM written. Like it or not, people can recognize LLM-generated text pretty easily https://news.ycombinator.com/item?id=45868782

I am pretty much certain that parts of it were LLM-written, yes. This doesn't imply that the entire blog post is LLM-generated. If you're a good Bayesian and object to my use of "100%" feel free to pretend that I said something like "95%" instead. I cannot rule out possibilities like, for example, a human deliberately writing in the style of an LLM to trick people, or a human who uses LLMs so frequently that their writing style has become very close to LLM writing (something I mentioned as a possibility in an earlier reply; for various reasons, including the uneven distribution of the LLM-isms, I think that's unlikely here).

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#64
post #46

Earlier quoted context omitted.

> this parallel construction “not A, not B, not C” and “not A, not B, but C” are extremely common constructions in general. So common in fact that you did it in this exact reply. “This doesn't mean that the entire article was written by an LLM, nor does it mean that there's not useful information in it. Regardless, given the amount of low effort LLM-generated spam that makes it onto HN, I think it is fairly defensibl…

Blog spam doesn’t intersperse the drivel with literary narrative beats and subsection titles that sound like sci-fi novels. The greasy mixture of superficially polished but substantively vacuous is much more pronounced in LLM output than even the most egregious human-generated content marketing, in part because the cognitive entity in the latter case is either too smart, or too stupid, to leave such a starkly evident…

Is… is this from an LLM? Because this is the first time I’ve felt confident identifying text as no-human-writes-this-way.

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#65
post #55

Earlier quoted context omitted.

On the other fork where I responded to your claims with a direct and detailed response, you insisted that my comment “isn't really that interesting” and just disengaged. I’m not going to write another detailed explanation of why your “slop === AI” premise is flawed. Go reread the other fork if you’ve decided you’re interested. > I find it interesting that you believe this claim is wildly conspirational I don’t believ…

Human experts can reliably detect some kinds of long-form, AI-generated text using exactly the same sorts of cues I've outlined: https://arxiv.org/html/2501.15654v1 . You may take issue with the quality of the paper, but there have been very few studies like this and this one found an extremely strong effect. I am making an even more limited claim than the article, which is only that it's possible for "experts" (i.e.…

That’s a really interesting study. Thanks for sharing that.

This seems like the kind of thing to share when making a bold claim about being able to detect AI with high confidence. This is a lot more weighty than not so subtly asserting that I’m too dumb to recognize AI.

> a human deliberately trying to sound like an LLM (I feel the need to reiterate this as many of the counterarguments to what I'm saying are to claims of this form).

I assume this is a reference to me. To be clear, I was never referring to humans specifically attempting to sound like AI. I was saying that a lot of formulaic stuff people attribute to AI is simply following the same patterns humans started, and while it might be slop, it’s not necessarily AI slop. Hence the AITA rage bait example.

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#66
post #65

Earlier quoted context omitted.

Human experts can reliably detect some kinds of long-form, AI-generated text using exactly the same sorts of cues I've outlined: https://arxiv.org/html/2501.15654v1 . You may take issue with the quality of the paper, but there have been very few studies like this and this one found an extremely strong effect. I am making an even more limited claim than the article, which is only that it's possible for "experts" (i.e.…

That’s a really interesting study. Thanks for sharing that. This seems like the kind of thing to share when making a bold claim about being able to detect AI with high confidence. This is a lot more weighty than not so subtly asserting that I’m too dumb to recognize AI. > a human deliberately trying to sound like an LLM (I feel the need to reiterate this as many of the counterarguments to what I'm saying are to claim…

Thanks for engaging thoughtfully! FWIW I actually looked this article up because I was interested in your claim that even experts couldn't perform these tasks, something I hadn't heard before--I'm not actually ignoring what you're saying. It's actually very nice to have a productive conversation on HN :)

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#67
post #38
post #27

Earlier quoted context omitted.

I feel like it’s on every other article now. The “this is ai” comments detract way more from the conversation than whatever supposed ai content is actually in the article. These ai hunters are like the transvestigators who are certain they can always tell who’s trans.

No. These articles are annoying to read, the same dumb patterns and structures over and over again in every one. It's a waste of time; the content gives off a generic tone and it's not interesting.

So is the vast majority of comments on HN (and in any comment section of any website) well before LLMs came into being, yet we give them a benefit of doubt. Users on forums tend to behave in a starkly bot-like way, often having a very limited set of responses pertaining to their particular hobby horses, so much so that others could easily predict how the most prolific users would react to any topic and in what precise words.

Now, apparently, we have a generation of "this is AI slop!" "bots".

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#68
post #64

Earlier quoted context omitted.

Blog spam doesn’t intersperse the drivel with literary narrative beats and subsection titles that sound like sci-fi novels. The greasy mixture of superficially polished but substantively vacuous is much more pronounced in LLM output than even the most egregious human-generated content marketing, in part because the cognitive entity in the latter case is either too smart, or too stupid, to leave such a starkly evident…

Is… is this from an LLM? Because this is the first time I’ve felt confident identifying text as no-human-writes-this-way.

I don’t usually speak like this on Hacker News, but for fucks sake, just give it a fucking rest already, you utter pillock.

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#69
post #34

So they turned on GC after every allocate ("GC stress"), and "With GC.stress = true, the GC runs after every possible allocation. That causes immediate segfaults because objects get freed before Ruby can even allocate new objects in their memory slots." That would seem to indicate a situation so broken that you can't expect anything to work reliably. The wrong-value situation would seem to be a subset of a bigger pro…

That’s exactly what it was. He discovered the customer was using a version of ffi that had this “use-after-free” (ish) bug, but the question “is this actually what my customer was seeing or is there _another_ bug lurking” still needed to be answered.

It's nice that there is only a few weird behaviors produced. Often use-after-free leads to so many different random bugs, you might gorble a hubalu.

Re: When your hash becomes a string: Hunting Ruby's million-to-one memory bug

#70
post #62

Earlier quoted context omitted.

> believing that you (and whatever group you identify as “us”) are able to deterministically identify AI text I think you will find the OP said no such thing. They instead said they identified a mixture of writing styles consistent with a human author and an LLM. The OP says nothing about deterministically identifying LLMs, only that the style of specific sections is consistent with LLMs leading to the conclusion.

I think you find OP absolutely did say that. > Parts of it were 100% LLM written. Like it or not, people can recognize LLM-generated text pretty easily https://news.ycombinator.com/item?id=45868782

Thanks for adding the quote, that is a different part of the post than I was focusing on.

I still think that's a far cry from deterministically recognizing LLM-generated text. At least the way I would understand that would be an algorithmic test with very low rates of both false positives and false negatives. Instead I understood the OP to be saying that people have an intuitive sense of LLM generated text with a relatively low false negative rate.

I am certain that the skill varies widely between individuals, but in principle there is no reason to suspect that with training humans could not become quite good at recognizing low effort (no attempt at altering style) LLM generated content from the major models. In principle it is no different than authorship analysis used in digital forensics, a field that shows fairly high accuracy under similar conditions.

Post reply on HN