Live data from Hacker News

Opus 4.7 knows the real Kelsey

theargumentmag.com

241–250 of 284 posts

Re: Opus 4.7 knows the real Kelsey

#241
post #164

Earlier quoted context omitted.

How do you know, how the model works? If there was an index of all Micken's writings, or even if the model searched the web before feeding the response to you, you wouldn't know by observing from the outside.

i suppose a quick test would be getting the model to write down Micken's essay end to end. if the original essay was stuffed within the prompt window. the result will be word accurate. unless this is a model trained specifically on Micken's essay (which claude is not).

This seems like a classic case of doing it being proof that it can happen, but not doing it being insufficient proof that it's impossible. I don't think there's a "quick test" of whether there might be a more effective prompt that would cause it to reproduce more effectively.

Re: Opus 4.7 knows the real Kelsey

#242
post #192

More people should have been aware that human text contains a lot of identifiable information, and a dumb statistical model could do this a decade ago. (There were show hns with Hn user similarity analysis that used a deceptively simple model (if I remember it used like most likely word pairs only) and it was very effective. It got taken down, but the cat has always been out the bag). So your "anonymous" account coul…

Growing up on MUDs, people could clock someone on a completely different, graphical game from their text patterns.

Re: Opus 4.7 knows the real Kelsey

#243
post #123
post #121

It's funny: publishing work offline in books and magazines is perhaps more anonymous in the age of AI. I pasted in a number of passages from books on my bookshelf. Predictably, stuff that I read for my English degree in university is largely in the training data and easily identifiable. Stuff from regional authors or is slightly adjacent to the cultural mainstream makes no impression.

To clarify, because a number of posts here sort of suggest the confusion: the article here isn't about the LLM recognizing works that were in the training data. EG, The Old Man and the Sea off the shelf. It's about pegging the author of novel texts, like, say, some letter written by Hemmingway that gets discovered next week and was never before digitized.

Yes, that makes sense. However, unless there's a significant corpus of an author in the training data it won't recognize them. One of the author's that I fed into Claude was a passage from the book Leepike Ridge by ND Wilson. Wilson has written online and in print quite a bit, but Claude couldn't guess the author and guessed that it was a passage from a noir crime novel.

Wilson is a fairly idiosyncratic writer with a distinct style, yet even still Claude couldn't guess correctly from a currently published book.

I suspect that what's going on here (like other's are suggesting in this thread) is that Claude is in some way biased towards certain sets of authors by its training.

Re: Opus 4.7 knows the real Kelsey

#244

The joke's on you all for willingly posting this content online for it to later be harvested by AI. Nobody is forcing you to use these systems. The hackers have always said this moment, or something like it, would come, from beneath their canopies of tin foil. I've posted almost nothing online - not under pseudonyms nor real names - for over a decade. I sat on this HN username for almost 12 years before making a sing…

How do you propose a journalist work without posting their writing online?

Journalists by definition cannot be anonymous. That's why its a dangerous job.

Re: Opus 4.7 knows the real Kelsey

#245
post #234
post #138

Earlier quoted context omitted.

> it correctly identified it as an imitation of James Mickens How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.

I fed an unpublished draft of mine to an AI. I saw it searching the internet and prompted it with the fact it could stop searching, it was not published. From there it guessed that it was me on the spot, which I thought was kind of funny. Can't deny the meta-logic there. It referred to me by my login name on the AI site rather than the name it would have used if it actually found my website, so I think it was more lo…

The author specifically discusses their efforts to avoid this sort of information leak which would obviously poison the result.

Re: Opus 4.7 knows the real Kelsey

#246
I am extremely skeptical of any of these claims, and of other commenters saying they replicated this.

First, the author fed an unpublished draft to Anthropic's hosted model. I assume they did this from their personal account, that may include a credit card or at the very least a pseudonymous name that is uniquely identifiable.

Then, the author fed an unpublished draft to Anthropic's hosted model, except in Incognito or whatever. We are led to assume that, whatever the author did for the second submission, they did so in a way so that Anthropic could not correlate both distinct requests from one another. Perhaps on a second subscription? They don't say. I am highly skeptical they airgapped their requests properly so that it doesn't look like the same user is making the request to the same hosted model.

Then, the author asked a friend to publish the draft. A friend, of which there is probably a digital trail that maps the relationship of the author to their friend.

All of this metadata could be crunched on the backend before the black box spits out a response.

Across all these datapoints, I have high confidence a model of this caliber could put two and two together and determine that the author penned the drafts, not solely because of stylometry, but because there is a clear behavioral pattern tying all three events together.

An assumption made here is that Anthropic doesn't train on chats. Though the author opted out of training on their chats, and session memory, how could you trust a hosted model to respect such opt outs?

Re: Opus 4.7 knows the real Kelsey

#247
I am an absolute nobody on the internet, so it didn't get me at all. But the names it did guess were some good bloggers, some which I had yet to follow.

So do this even if it has no chance of getting it right for you.

Re: Opus 4.7 knows the real Kelsey

#249

I am extremely skeptical of any of these claims, and of other commenters saying they replicated this. First, the author fed an unpublished draft to Anthropic's hosted model. I assume they did this from their personal account, that may include a credit card or at the very least a pseudonymous name that is uniquely identifiable. Then, the author fed an unpublished draft to Anthropic's hosted model, except in Incognito…

How do you explain other people on these chats making similar claims? Everybody is making the same mistakes?

Re: Opus 4.7 knows the real Kelsey

#250

I am extremely skeptical of any of these claims, and of other commenters saying they replicated this. First, the author fed an unpublished draft to Anthropic's hosted model. I assume they did this from their personal account, that may include a credit card or at the very least a pseudonymous name that is uniquely identifiable. Then, the author fed an unpublished draft to Anthropic's hosted model, except in Incognito…

How do you explain other people on these chats making similar claims? Everybody is making the same mistakes?

It's explained by the near impossibility of isolating requests from each other, and chain of custody of divulged information.

If I send a prompt from identity A, which is the true user identity, you have possibly sent all of identity A metadata to be ingested alongside the prompt to generate response X.

If I /then/ send the prompt from identity B, the prompt has been answered before with metadata from identity A. The black box can consult metadata from response X to generate response Y, thus possibly correlating response Y with the prompt sent by identity A.

Post reply on HN