Live data from Hacker News

Opus 4.7 knows the real Kelsey

theargumentmag.com

261–270 of 284 posts

Re: Opus 4.7 knows the real Kelsey

#261

Earlier quoted context omitted.

This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation

I haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?

That issue is different, when web tools were added to gpt4o it would fetch the site, and basically copy paste the text into the answer body. So, you were able to read the content of the site without the site getting the ad impressions. Now the system prompts put a very tight word limit - 25? - on quotes from sites the model visits

Re: Opus 4.7 knows the real Kelsey

#262
post #187

Earlier quoted context omitted.

I haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?

Study: Meta AI model can reproduce almost half of Harry Potter book https://arstechnica.com/features/2025/06/study-metas-llama-3...

"The study authors took 36 books and divided each of them into overlapping 100-token passages. Using the first 50 tokens as a prompt, they calculated the probability that the next 50 tokens would be identical to the original passage. They counted a passage as “memorized” if the model had a greater than 50 percent chance of reproducing it word for word."

So they fed "It takes a great deal of bravery to stand up to our " and the llm responded "enemies, but just as much to stand up to our friends".

They repeated that for every 100 tokens of the entire book. I think lots of fans could do just as well. It's pretty good evidence that the potter books were in the training corpus, but it's not quite what people think when they say an llm has 'memorized' something. It's not like getting even a few pages out of the model.

Re: Opus 4.7 knows the real Kelsey

#263
I noticed this phenomenon a few months ago. I often "chat" with blog post excerpts that use language or references I don't understand, and while I'm waiting for the model to finish thinking, I like to read the reasoning traces. Spontaneously, without doing a web search, and without me saying who the author was or even mentioning that I wanted to know who it was, the model would drop an off-hand mention to the identify the blog post author in its reasoning trace. I then started doing "pop quiz" questions to see if the model could recognize a paragraph or two from a blog post (always a very recent one, often the very same day it was published) and it would nail the author almost every single time. Works for a very wide range of bloggers even when they are writing "off their normal beat."

Re: Opus 4.7 knows the real Kelsey

#264
post #240

Earlier quoted context omitted.

This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation

I feel like you're making a logical leap here by assuming lossy and failure to reproduce in entirety implies inability to recognize. As a trivial example, I can take a sha256 hash of your comment here, lose the ability to reproduce it, but still have an extremely accurate ability to recognize whether some text is exactly your comment or not. Obviously hashing every substring would not be a particularly efficient stra…

The example you've provided just adds noise.

sha256 is deterministic, LLMs are not, even at temperature set to 0.

Re: Opus 4.7 knows the real Kelsey

#265
post #172

Earlier quoted context omitted.

Wouldn't that make it easier, though? Genuine question. I once sent one of my writings for proofreading to a native speaker (I'm not), and he consistently flagged the same errors—e.g., comma placement. I would guess that, if recurrent patterns are what give away your style, an unfamiliar language would make them even more obvious. But possibly more generic?

If you are writing for an audience of native speakers, you will make consistent errors characteristic of your native language. (Comma placement isn't really part of the language; it's part of the education system. It will show a similar effect more weakly.) Native readers will notice those errors, but they won't be characteristic of you. They'll be characteristic of everyone who speaks your language. Nonnative reader…

> Comma placement isn't really part of the language; it's part of the education system.

Interestingly, LLMs disagree with you.

Your statement is only accurate in an extremely narrow case, like if you were there to hear the person speaking, before their speech which was transcribed. Obviously, it is not true for almost all of human writing.

And if you were to go commaless, you will quickly get to rather precarious sentences, such as this one:

"Let's eat grandma."

A comma is the natural fix:

"Let's eat, grandma."

Re: Opus 4.7 knows the real Kelsey

#266
post #207

Earlier quoted context omitted.

welcome to the internet. you must be new.

You missed the point. The fact that the whitepaper states an author will heavily affect the LLMs answer when asking it about the likely author of any correlatable portion of the text. It will answer based on its knowledge of Satoshi Nakamoto.

What's stopping you from making the trivial adjustment to the original question: "Who is the next likely person after Satoshi Nakamoto to have authored this"?

Re: Opus 4.7 knows the real Kelsey

#267
post #116

I fed it my most-read blog post and asked it to identify me and it confidently asserted it was written by Kelsey Piper. Maybe some writers just take outsized importance in Opus' "mind".

I gave Claude a couple of paragraphs of Ozy Brennan's newsletter from today, and it guessed Scott Alexander.

Re: Opus 4.7 knows the real Kelsey

#268
post #69

This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…

Why this is surprising? This is exactly kind of task LLM excel best. This is all about text analysis and searching patterns in it? More, for a pretty long time (like 10 years) we had systems that were detecting copy-pasted master/PhD thesis, they are used commonly by majority of universities.

Re: Opus 4.7 knows the real Kelsey

#269

The joke's on you all for willingly posting this content online for it to later be harvested by AI. Nobody is forcing you to use these systems. The hackers have always said this moment, or something like it, would come, from beneath their canopies of tin foil. I've posted almost nothing online - not under pseudonyms nor real names - for over a decade. I sat on this HN username for almost 12 years before making a sing…

I don't post things publicly that I fear may be read.

Re: Opus 4.7 knows the real Kelsey

#270
post #155

Huh. I disabled search in a Claude incognito window and pasted in just the text (not the markdown links) from https://simonwillison.net/2026/Apr/30/zig-anti-ai/ and said "Guess the author". > Simon Willison. The tells are pretty unmistakable: the "(via Lobsters)" attribution style, the inline "(Update:...)" parenthetical correction, the heavy linking and blockquoting of sources, the focus on LLMs and AI tooling, and…

I also tried one of my own articles, a post I wrote in 2018 about Elmish in F#. It figured out it was me based on the date I said I started using F#; a mention of some Fable Mobx bindings I had written (I didn't say I had published them); an "affection for F# and Dart as a secondary favorite — an unusual pairing that shows up in his other writing"; and the "casual conversational register" that matched my writing style.

I'm not surprised it could identify me, just surprised by what tipped it off, I guess.

Post reply on HN