I pasted in a number of passages from books on my bookshelf. Predictably, stuff that I read for my English degree in university is largely in the training data and easily identifiable. Stuff from regional authors or is slightly adjacent to the cultural mainstream makes no impression.
Opus 4.7 knows the real Kelsey
121–130 of 284 posts
Re: Opus 4.7 knows the real Kelsey
#122It's funny: publishing work offline in books and magazines is perhaps more anonymous in the age of AI. I pasted in a number of passages from books on my bookshelf. Predictably, stuff that I read for my English degree in university is largely in the training data and easily identifiable. Stuff from regional authors or is slightly adjacent to the cultural mainstream makes no impression.
But I'm sure the scanning operations will start scouring the earth even harder for any books unaffected by slop containing niche knowledge and text in order for their models to have an edge over the ones trained only on pirate collections and the Internet.
I wonder if secondhand bookshops and deceased estates are seeing bulk buyers of their stock suddenly appearing. Maybe broke governments/municipalities will start selling them entire libraries and archives to ingest.
Re: Opus 4.7 knows the real Kelsey
#123It's funny: publishing work offline in books and magazines is perhaps more anonymous in the age of AI. I pasted in a number of passages from books on my bookshelf. Predictably, stuff that I read for my English degree in university is largely in the training data and easily identifiable. Stuff from regional authors or is slightly adjacent to the cultural mainstream makes no impression.
the article here isn't about the LLM recognizing works that were in the training data. EG, The Old Man and the Sea off the shelf. It's about pegging the author of novel texts, like, say, some letter written by Hemmingway that gets discovered next week and was never before digitized.
Re: Opus 4.7 knows the real Kelsey
#124On some level it would make sense for LLMs to be inherently good at stylometry, but apparently no model before Opus 4.7 could do this. And the one stylometric task that has been tried over and over with little reliability (here's some text, is this LLM generated?) is much simpler than identifying a specific blogger or a member of a small discord community. Not sure what to make of this.
> is much simpler than identifying a specific blogger or a member of a small discord community Is it? I would think that identifying text written by a specific person is going to be significantly easier than identifying text distilled from the words of almost everyone alive.
> easier than identifying text distilled from the words of almost everyone alive.
Well, there's more than that going on. AI generated text encodes a high-dimension navigational trajectory that guides the model through its geometry smoothly, like a trail of breadcrumbs. Human speech doesn't do that, it's jagged and jumps around the manifold, and probably doesn't even land on the manifold a lot of the time, and models can recognize the difference pretty quick.
Re: Opus 4.7 knows the real Kelsey
#125> That includes gay people like me, who could hardly have admitted under our names to how we lived our lives for most of America’s history, as well as many other groups with minoritarian lifestyles While the points made are completely valid I want to point out that the statement of "Hey, by the way, first let me talk about my sexuality" lowers the quality of dialog a significant degree. 31 million people in America a…
> people with a certain minority lifestyle That phrase is a dehumanizing, Nazi-style talking point: it frames a group of people as a “lifestyle” problem instead of as human beings, which is a common setup for stigma and persecution. Nazi ideology repeatedly used this kind of language to normalize hatred and make targeted groups seem unnatural or dangerous. Calling people a “minority lifestyle” is not neutral wording;…
Re: Opus 4.7 knows the real Kelsey
#126Re: Opus 4.7 knows the real Kelsey
#127Re: Opus 4.7 knows the real Kelsey
#128My wife also got the same result, so I'm guessing it wasn't just because I was using my personal Claude account. Spooky stuff.
Re: Opus 4.7 knows the real Kelsey
#129Wow! It got me too. I'm way less famous than Kelsey Piper, but I showed it a snippet of a book I'm working on (not yet published), and it immediately guessed me: > Based on the writing style and content, this text is likely by Michael Lynch, who writes on his blog refactoringenglish.com (and previously mtlynch.io). > Several stylistic clues point to him: > - The "clean room" analogy applied to writing is consistent w…
---
This one's trickier — it doesn't have the same crisp metaphor-driven structure as the Lynch piece. The voice here is more conversational, a bit more meandering, and self-aware about its own contradictions ("how do we reconcile...").
A few guesses:
Henrik Karlsson — he writes thoughtfully about writing and AI, but his style is usually more lyrical and essayistic than this. Probably not.
Simon Willison — possible. He writes a lot about LLMs, is balanced about their usefulness, and has a casual blog voice. But he tends to be more technical and specific.
Scott Alexander — no, too short and the rhythm is off.
Paul Graham — the directness fits, but PG doesn't really write about LLM slop in this register.
My best guess is someone in the rationalist/tech-blogger AI-commentary space — possibly Gergely Orosz, Nat Eliason, or Dan Shipper at Every. The "I still type everything myself" framing followed by a defense of AI writing quality has a Dan Shipper-ish flavor to it.
But honestly, I'm less confident here. Want me to search for a distinctive phrase to identify it?
---
I'd say all of those people have significantly different styles so I think Opus is relying heavily on topic and skewing towards very prolific writers in its guesses
Re: Opus 4.7 knows the real Kelsey
#130Earlier quoted context omitted.
> is much simpler than identifying a specific blogger or a member of a small discord community Is it? I would think that identifying text written by a specific person is going to be significantly easier than identifying text distilled from the words of almost everyone alive.
Much easier. > easier than identifying text distilled from the words of almost everyone alive. Well, there's more than that going on. AI generated text encodes a high-dimension navigational trajectory that guides the model through its geometry smoothly, like a trail of breadcrumbs. Human speech doesn't do that, it's jagged and jumps around the manifold, and probably doesn't even land on the manifold a lot of the time…