Earlier quoted context omitted.
This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation
I haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?
Opus 4.7 knows the real Kelsey
261–270 of 284 posts
Re: Opus 4.7 knows the real Kelsey
#262Earlier quoted context omitted.
I haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?
Study: Meta AI model can reproduce almost half of Harry Potter book https://arstechnica.com/features/2025/06/study-metas-llama-3...
So they fed "It takes a great deal of bravery to stand up to our " and the llm responded "enemies, but just as much to stand up to our friends".
They repeated that for every 100 tokens of the entire book. I think lots of fans could do just as well. It's pretty good evidence that the potter books were in the training corpus, but it's not quite what people think when they say an llm has 'memorized' something. It's not like getting even a few pages out of the model.
Re: Opus 4.7 knows the real Kelsey
#263Re: Opus 4.7 knows the real Kelsey
#264Earlier quoted context omitted.
This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation
I feel like you're making a logical leap here by assuming lossy and failure to reproduce in entirety implies inability to recognize. As a trivial example, I can take a sha256 hash of your comment here, lose the ability to reproduce it, but still have an extremely accurate ability to recognize whether some text is exactly your comment or not. Obviously hashing every substring would not be a particularly efficient stra…
sha256 is deterministic, LLMs are not, even at temperature set to 0.
Re: Opus 4.7 knows the real Kelsey
#265Earlier quoted context omitted.
Wouldn't that make it easier, though? Genuine question. I once sent one of my writings for proofreading to a native speaker (I'm not), and he consistently flagged the same errors—e.g., comma placement. I would guess that, if recurrent patterns are what give away your style, an unfamiliar language would make them even more obvious. But possibly more generic?
If you are writing for an audience of native speakers, you will make consistent errors characteristic of your native language. (Comma placement isn't really part of the language; it's part of the education system. It will show a similar effect more weakly.) Native readers will notice those errors, but they won't be characteristic of you. They'll be characteristic of everyone who speaks your language. Nonnative reader…
Interestingly, LLMs disagree with you.
Your statement is only accurate in an extremely narrow case, like if you were there to hear the person speaking, before their speech which was transcribed. Obviously, it is not true for almost all of human writing.
And if you were to go commaless, you will quickly get to rather precarious sentences, such as this one:
"Let's eat grandma."
A comma is the natural fix:
"Let's eat, grandma."
Re: Opus 4.7 knows the real Kelsey
#266Earlier quoted context omitted.
welcome to the internet. you must be new.
You missed the point. The fact that the whitepaper states an author will heavily affect the LLMs answer when asking it about the likely author of any correlatable portion of the text. It will answer based on its knowledge of Satoshi Nakamoto.
Re: Opus 4.7 knows the real Kelsey
#267I fed it my most-read blog post and asked it to identify me and it confidently asserted it was written by Kelsey Piper. Maybe some writers just take outsized importance in Opus' "mind".
Re: Opus 4.7 knows the real Kelsey
#268This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…
Re: Opus 4.7 knows the real Kelsey
#269The joke's on you all for willingly posting this content online for it to later be harvested by AI. Nobody is forcing you to use these systems. The hackers have always said this moment, or something like it, would come, from beneath their canopies of tin foil. I've posted almost nothing online - not under pseudonyms nor real names - for over a decade. I sat on this HN username for almost 12 years before making a sing…
Re: Opus 4.7 knows the real Kelsey
#270Huh. I disabled search in a Claude incognito window and pasted in just the text (not the markdown links) from https://simonwillison.net/2026/Apr/30/zig-anti-ai/ and said "Guess the author". > Simon Willison. The tells are pretty unmistakable: the "(via Lobsters)" attribution style, the inline "(Update:...)" parenthetical correction, the heavy linking and blockquoting of sources, the focus on LLMs and AI tooling, and…
I'm not surprised it could identify me, just surprised by what tipped it off, I guess.