Live data from Hacker News

The revolt of the reader

bcantrill.dtrace.org

51–60 of 311 posts

Re: The revolt of the reader

#51
Hi Bryan,

I liked your piece, and agree with almost all of it, but I'm surprised by your faith in the accuracy of Pangram at detecting AI writing. Is your faith based on testing it with lots of writing of known origins, or are you just saying that it reaches the same conclusion that you do as a talented human?

In particular, I wondered if you have tried running all of your own writings through it to verify that it thinks you are human. I was struck by Freddie deBoer's recent piece where he did this and said it often failed: https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...

What percentage of false positive rejections would you find acceptable? Would you accept this even if it forced you to change the way you write?

Re: The revolt of the reader

#52
post #35

I would love to use Pangram but they simply don’t allow signing up with my custom email domain. The error was “This email address can't be used for signup. Please use a different email.” I’m not about to create a Gmail is to use your service. To me the attack on the decentralized nature on Internet infrastructure is no less serious than the attack on the human provenance of writing itself.

Thanks for writing about that. I appreciate you taking the time to call out a bad actor like that.

Re: The revolt of the reader

#53
post #41
post #26

Earlier quoted context omitted.

I don't think this is about edge cases where someone has successfully disguised the writing to some degree: the current crop of LLMs have some pretty blatant (and frankly annoying) habits by default, ones that are hard to miss once you have read a decent amount of their output. If I had to describe them broadly, I would say they are a collection of habits which are common in certain kinds of persuasive and emotive wr…

> the current crop of LLMs have some pretty blatant This is just "em dash redux." Except now we've moved on to accusing anyone who does "It's not X. It's Y." of being AI. In six months, it'll be "use of the word 'petrichor'" or something.

I do think there is a tendency to over-index on one or two particularly straightforward tells, and for any given feature of LLM writing you can find places where people do also use that feature (they had to learn it from somewhere, and in a lot of cases it is good writing practice — for the context in which it is used). But I'm not talking about just that, but also the general tone issue: it's bad writing regardless because it's in most cases just not appropriate for the context it's been written in.

(TBH I think the biggest likelihood for false positives comes from heavy LLM users picking up their tics: it's a natural tendency and I've already seen a few cases where it seems like that has happened).

Re: The revolt of the reader

#54
post #41
post #26

Earlier quoted context omitted.

I don't think this is about edge cases where someone has successfully disguised the writing to some degree: the current crop of LLMs have some pretty blatant (and frankly annoying) habits by default, ones that are hard to miss once you have read a decent amount of their output. If I had to describe them broadly, I would say they are a collection of habits which are common in certain kinds of persuasive and emotive wr…

> the current crop of LLMs have some pretty blatant This is just "em dash redux." Except now we've moved on to accusing anyone who does "It's not X. It's Y." of being AI. In six months, it'll be "use of the word 'petrichor'" or something.

Idk, it's more like "your writing is cliché and I don't feel like reading it because I've already read something that sounded similar countless times and it wasn't worth the read". The source of the clichés being an LLM. And maybe now humans are writing the same way as LLM output, I still am not going to read all that, sorry. If I see a sea of clichés, I'm going the other way.

I'm also not reading pumpkin spice murder mysteries for a similar reason. I'm also not reading stories where everybody clapped. Actually, I'm already familiar with petrichor, so unless someone has surrounded the word "petrichor" with non-cliché prose, I'm also not going to read all that.

Re: The revolt of the reader

#55

This meme of trying to make it sound like LLM text is so obvious is a joke. It’s literally not, you can tell it to write in literally any style and given just a bit of an example of a person’s writing style, frontier models copy it completely and effectively. This argument can probably be leveled at vanilla raw output from an LLM, but even the slightest attempt at obfuscation bears solid fruit.

Well, give it a shot -- you'll likely find that that technique doesn't work nearly as well (at least with Pangram 4) as you think it might. When we had Max on the podcast[0], Adam explicitly asked him about exactly this (after all, you can give an LLM access to Pangram and let it iterate!), and Max reported that someone had attempted to do this -- and ended up burning through $700 in tokens and had a "sad Claude." An…

99% of college essays and pretty much everything “product” in corporate America is now LLM generated with some marginal oversight. It passes muster for the most part.

Re: The revolt of the reader

#56
post #51

Hi Bryan, I liked your piece, and agree with almost all of it, but I'm surprised by your faith in the accuracy of Pangram at detecting AI writing. Is your faith based on testing it with lots of writing of known origins, or are you just saying that it reaches the same conclusion that you do as a talented human? In particular, I wondered if you have tried running all of your own writings through it to verify that it th…

My experience is using Pangram quite often with lots of writing of all flavors (including a bunch of known origin).

As for my own writing, I didn't do this experiment, but one of my co-workers did -- and over 176 posts spanning 22 years, all 176 (well, 177 now with my latest) are 100% human. This is not hugely surprising in that (in addition to me having actually written them!) my voice is very... distinctive. What would be more entertaining would be to try to get an LLM to write like me and fool Pangram that way. I still think that this would be difficult based on the experiences that I've heard, but it wouldn't surprise me if you could pull it off (and I would assuredly find the result entertaining!).

In the dimensions that we use Pangram in the most actionable sense (namely, to audit our own public writing), I am unconcerned about false positives, and leave it to Oxide authors to rework/recast as needed. (Though it sounds like Freddie didn't even need to do that -- he just needed to provide a longer sample.)

Re: The revolt of the reader

#57
post #24

Earlier quoted context omitted.

“Not every one of those 11,970 carries the same weight, and the paper is careful about that rather than rounding it up.” AAUGH IT BURNS

Right?! When I hit "That work was genuinely valuable" I literally hollered in exasperation.

For me it was “the results speak for themselves” and then simply a (large) number of automated tests run that never had human eyes.

Yes, quantity famously has a quality all its own, but perhaps not where correctness checks for something this central is concerned.

Re: The revolt of the reader

#58
For me, writing is an activity of expressing my feelings and conveying my thoughts. I rarely left that to LLM simply because one does not contract out activities one cherishes.

I (am kinda forced to) use LLM to generate maybe 40% of the code at work, that is after my review and modifications. But I pretty much wrote all of the comments by myself. I can get into the flow by writing comments.

Re: The revolt of the reader

#59
post #40

This meme of trying to make it sound like LLM text is so obvious is a joke. It’s literally not, you can tell it to write in literally any style and given just a bit of an example of a person’s writing style, frontier models copy it completely and effectively. This argument can probably be leveled at vanilla raw output from an LLM, but even the slightest attempt at obfuscation bears solid fruit.

This isn't very effective on any models released in recent years. With older ones, you used to be able to influence writing style significantly by just putting examples in the context, but newer models have gone through so much assistant RLHF, they really want to revert back to their default "assistant voice" during their turn. You can still influence their writing style in a broad manner that might look correct at a…

I think nowadays defeating the detection probably looks like finetuning a smaller LLM and getting it to paraphrase the text from the other one (or just using a more obscure finetune: it'll probably have its own cliches and habits but it will be at least different). As an added bonus this also likely removes the fingerprinting from the output as well. But I think most people are not going to bother with this.

Re: The revolt of the reader

#60
post #38

> To those who read broadly, the hand of the LLM is so clear it’s as if the writer’s intellectual fly is open I dunno, man, according to Hardcover, I've read 76 fiction books this year, and I can't tell. All the "AI tells" fail the vibe check. I'm a writer and I get flagged by many of them. And according to PhD linguists with expertise in the field, most AI tells are just the equivalent of old wives' tales. https://w…

I think anyone claiming 100% accuracy is wrong, but the recent Claude models, for example, have a writing style that is sufficiently distinct that claiming people can't recognize it is like claiming you can't recognize the styles of particular famous authors. Yes, particular elements of their writing are going to be used by others, and it's possible to disguise their style or emulate it deliberately, but it's pretty hard to accidentally write like them.
Post reply on HN