This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…
That's neat, though it impresses me less that the article. Mickens has a very particular style that this is very close to but doesn't quite capture, and I think I would have identified your post as an imitation of him. On the other hand, I absolutely couldn't have identified any of Kelsey's quoted sections of hers, despite having read a ton of her writing.
Opus 4.7 knows the real Kelsey
231–240 of 284 posts
Re: Opus 4.7 knows the real Kelsey
#232Re: Opus 4.7 knows the real Kelsey
#233Can't wait to have to exchange stylometric encoders with my loved ones so that we can exchange truly private messages without losing our human touch.
intriguing, care to explain a bit more?
Re: Opus 4.7 knows the real Kelsey
#234This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…
> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.
It referred to me by my login name on the AI site rather than the name it would have used if it actually found my website, so I think it was more logic than an actual identification, but it had clearly corrupted the search enough to no longer be a valid test.
Which does make me wonder about the original article; if the AI has in context any sort of clue that the user is "Kelsey Piper" (a memory of their name, a username of kpiper or kelseyp, etc.), that will radically tip the balance in favor of the AI guessing that way just by the nature of LLMs. That is to say, it highly increases the odds of that guess even if it's wrong.
Even if that is the case, though, the general identifiability of writing remains true. It's been shown for a while with techniques a lot less powerful than a frontier LLM.
Re: Opus 4.7 knows the real Kelsey
#235But, yeah, I’m a nobody that has been blogging (very sporadically) publicly (and writing at length on forums like this one, with various handles loosely tied to my real identity) for twenty or more years (and by virtue of not trusting 3rd parties to host my content, most of it is actually still up) and Opus 4.6 (didn’t try 4.7) got me on the first try with just two paragraphs of an unpublished draft post (though it couldn’t come up with a convincing reason as to why it thought it was me).
Gemini and ChatGPT both clearly go off the subject matter rather than the stylistic clues; for the specific blog post I fed it which included mentions of “decoding” and “deciphering” and spoke of a tranche of legal documents (ok, it was the Epstein files, which I have been working on decoding), Gemini and ChatGPT both guessed “Molly White”, who seems to be a crypto-adjacent (currency not the real thing?) technical writer, and gave explanations that actually did explain why they arrived at that (wrong) answer.
So it seems Opus is indeed a bit special in this regard (and not limited to the latest 4.7 release)!
—-
What I would be more curious about is how well they can identify (open source) developers from their code. I’ve possibly publicly published more tokens in the form of OSS code than prose over the same period of time, in multiple languages and for completely different applications and environments. I’m sure there are style stylometric quirks associated with my coding style that persist across codebases, (though possibly somewhat stunted when contributing to others’ codebases to comply with the respective projects’ standards and styles) that should make it possible for an LLM that’s ingested code (and commits) to guess who’s who.
Edit:
Reading this self-same comment: I am apparently obsessed with parentheticals. Maybe my writing is more distinctive than I realized!
Re: Opus 4.7 knows the real Kelsey
#236Earlier quoted context omitted.
This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation
I haven't been following it well but isn't part of the NYT lawsuit against OpenAI that it sometimes spits out NYT articles verbatim?
There might be ten million people who have quoted Harry Potter at some point in their blogs or forum posts. There are only so many words in the books.
Re: Opus 4.7 knows the real Kelsey
#237This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…
That suggests it is picking up not only on style, but on the gap between authentic style and performed style. Useful for detecting pastiche, but pretty unsettling for pseudonymous writing.
Re: Opus 4.7 knows the real Kelsey
#238Re: Opus 4.7 knows the real Kelsey
#239This is blowing my mind. I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]: > Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a singl…
> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.
Re: Opus 4.7 knows the real Kelsey
#240Earlier quoted context omitted.
> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.
This is unlikely. The way model distribution works is that the model retains a lossy representation of James Micken's writing. Very likely, it cannot repeat Micken's writing verbatim. Neither can it reason about the training cutoff in this manner. It's a lossy representation