Live data from Hacker News

Opus 4.7 knows the real Kelsey

theargumentmag.com

151–160 of 284 posts

Re: Opus 4.7 knows the real Kelsey

#151
post #116

I fed it my most-read blog post and asked it to identify me and it confidently asserted it was written by Kelsey Piper. Maybe some writers just take outsized importance in Opus' "mind".

Yeah, style attribution is something big generalist models are usually pretty bad at, even on the material they have likely been trained on. Sure they are classifiers but this ability is limited, there's too much going on in them and they aren't magic. This needs a proper experiment, not anecdotal evidence.

Re: Opus 4.7 knows the real Kelsey

#152
If he does the same tests every time new models come out, and - I assume - uses the same dataset to do that, then is it not a possibility the said dataset is now part of the training set for the next round and therefore identifying who posted the text a fairly easy proposition ?

Re: Opus 4.7 knows the real Kelsey

#153

Earlier quoted context omitted.

> He explained that when he fed it snippets of the beginning of text, it would complete it in his voice and then sign it with his name. Is this public text already in the training set, or private text that might as well be written on the spot for the AI? I don't doubt AI can "fingerprint" you through your text (ideas, vocabulary, tone, etc), but those are different things, capability-wise

> I don't doubt AI can "fingerprint" you through your text (ideas, vocabulary, tone, etc), but those are different things, capability-wise The entire point of AI is pattern recognition, everything else is icing on the cake.

Imagine a cake with a truckload of icing on top.

Re: Opus 4.7 knows the real Kelsey

#155
Huh. I disabled search in a Claude incognito window and pasted in just the text (not the markdown links) from https://simonwillison.net/2026/Apr/30/zig-anti-ai/ and said "Guess the author".

> Simon Willison. The tells are pretty unmistakable: the "(via Lobsters)" attribution style, the inline "(Update:...)" parenthetical correction, the heavy linking and blockquoting of sources, the focus on LLMs and AI tooling, and the overall structure of an annotated link post commenting on someone else's writing. This reads exactly like a post from his blog at simonwillison.net.

Re: Opus 4.7 knows the real Kelsey

#156

Earlier quoted context omitted.

Problem is that it's been heavily contaminated with people speculating about who the author is. It would probably be difficult to get an unbiased answer out of it (although who knows - it's crazy that it can do this at all).

So train on pre 2009 mailing lost archive. Someone must be doing this surely.

Much better, train on the cypherpunk mailing list archive or anyone discussing e-cash on crypto forums or usenet from the 80's to the early 2010s

Re: Opus 4.7 knows the real Kelsey

#157
This ought to be guard-railed.

Doesn't seem like a valid use case for your average Joe to be able to identify anonymous authors at the click of a button.

Ofc state actors and proficient hackers can do most of it already, but this has genuine risk attached.

Re: Opus 4.7 knows the real Kelsey

#158

This ought to be guard-railed. Doesn't seem like a valid use case for your average Joe to be able to identify anonymous authors at the click of a button. Ofc state actors and proficient hackers can do most of it already, but this has genuine risk attached.

You have the vibes of people who think license plate numbers are private.

Re: Opus 4.7 knows the real Kelsey

#159
Stylometry has existed for decades, and there's no way an LLM is stronger at that job than a specialized piece of software (it's not more realistic than expecting Opus to beat Stockfish at chess).

In practice, you've never been anonymous while posting on the internet and AI isn't changing anything on that front. Or rather: if anything, AI can help you become more anonymous than before, since it can be used to hide your identity from stylometry by rewriting your prose before publishing.

Re: Opus 4.7 knows the real Kelsey

#160
post #138
post #69

This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…

> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.

This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner.

It's a lossy representation

Post reply on HN