Live data from Hacker News

Stealing Reasoning Traces from Proprietary LLM APIs

stolen-thoughts.com

291–300 of 325 posts

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#291
I really wonder how much of safeguarding with SOTA models is actually just "Don't do that" in a prompt?

I'm building agent to control machines at work and have to rely on small models, which are especially bad at remembering rules like this, so it is very apparent to me that soft rules like this are pretty much useless. Might be different for these huge models, but still it seems very shaky to me.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#292
post #285

Cool find, but can't help myself thinking that registering a domain name and submitting a paper on this to Arxiv is a bit... much. The content here could fit in a tweet or a short blog post as well. Not sure about the scientific novelty here as we're basically poking around the very top layers of someone else's software stack?

Growth hacking

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#293
post #85

Earlier quoted context omitted.

Correct, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches

For all of the years of research, thinking and talking about model alignment, safety, confinement, etc, when it comes down to it these companies appear to be entirely incompetent.

Alignment research was always, at best, security theatre.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#294

Earlier quoted context omitted.

For all of the years of research, thinking and talking about model alignment, safety, confinement, etc, when it comes down to it these companies appear to be entirely incompetent.

This isn't a safety issue, it's LLM companies trying to be opaque and stop distillation.

[dead]

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#295
Last year's reports questioning whether reasoning blocks were actually reflected in the final response were why I stopped using reasoning models altogether. I switched to a separate pipeline and have used that ever since. It's good to see that the concern didn't remain just a suspicion.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#296
post #254

Earlier quoted context omitted.

Yet we have the option to decide not to participate. That is an option.

I’m just driving by here but they bill by tokens — it’s a stretch to turn around and deny your right to see them. And it’s especially egregious when the tokens admittedly, routinely do the opposite of what you instructed. But personally it’s not about right and won’t it’s just blatant bullshit.

And lawyers bill by 6-minute increments, yet that doesn't mean you get access to all of a law firm's internal discussions and notes about you and your case.

Just because you paid for the lawyer time/LLM tokens doesn't mean you get access to everything that happened within that time/tokens.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#297
post #9
post #4

> We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, ... Ha! I've been wondering if replaying across models would work, ever since https://blog.cryptographyengineering.com/2026/05/29/fooling-... I'm honestly rather curious if this was intentionally allowed, it's the sort of validation that's easy to miss (particularly if you're wading into the vibe waters). Seem…

If you didn’t allow it, you wouldn’t be able to change models in the same conversation, as key parts of the context would be lost. Wouldn’t surprise me if the providers just remove that ability and lock the model once the conversation starts.

AFAIK no provider guarantees compatibility of reasoning traces, even in the same model generation, and we've in practice seen most of the big LLM APIs throw errors indicating incompatibility (at least transiently) when switching models. The only stable solution right now is to just throw away reasoning traces whenever a model is switched.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#298
post #227

Earlier quoted context omitted.

By some interpretations protecting your frontier model with strong safeguards from being distilled into an open source model without safeguards is a safety issue.

There's no way to make a model "safe", (whatever that means) since you can't know what users will do with the output. It's just PR. Closed models are also used for nefarious usage.

> There's no way to make a model "safe",

You could limit what it was training on in the first place - however that would damage capability - and it's difficult to curate the input, especially when the models can do 2+2. ie the choice is between model power and safety - and they choose power and everything else is a sticking plaster.

One thing I find amusing is the refusal of a lot of the models to now output a lab based protocol because of fears about 'weapons' - yet I can buy a textbook or simply read papers for exact protocols.

I find it hard to reason that a person who isn't motivated enough to read a paper or buy a book, is somehow enabled to make a biological weapon because of ChatGPT - despite them needed to buy a whole bunch of specialist equipment and reagents to do it.

Are there a whole bunch of proto-terrorists who are frustrated simply because they don't know where to start?

Maybe the only place their might be radicalized teenagers - but then that's perhaps a reason for keeping them off the internet full stop :-)

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#299

Earlier quoted context omitted.

There's no way to make a model "safe", (whatever that means) since you can't know what users will do with the output. It's just PR. Closed models are also used for nefarious usage.

> There's no way to make a model "safe", You could limit what it was training on in the first place - however that would damage capability - and it's difficult to curate the input, especially when the models can do 2+2. ie the choice is between model power and safety - and they choose power and everything else is a sticking plaster. One thing I find amusing is the refusal of a lot of the models to now output a lab ba…

That's also not possible, what's the worst problems enabled by LLM? Political propaganda, influence of population at scale, misleading advertising, social media bots... None of that will be filtered by a "safety" filter.

The knowledge to create weapons is already widespread, the idea that terrorists need chatgpt for that is laughable

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#300

Earlier quoted context omitted.

"copying is not theft/ stealing a thing leaves one less left/ copying it makes one thing more/ that's what copying's for"

Yes. Although, to be pedantic - stealing relocates (it doesn’t leave one less, it moves the only thing into another person’s possession), while copying duplicates. Copy vs move is IMHO accurate semantics.

i dont know what you're talking about: it's certainly one less for the victim! and if no theft occurred, nobody would be left with one less.
Post reply on HN