Live data from Hacker News

Stealing Reasoning Traces from Proprietary LLM APIs

stolen-thoughts.com

281–290 of 325 posts

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#281

Earlier quoted context omitted.

"Stealing" of non-rivalrous goods? I don't "think" it's not real. I know .

Reread what he called a morally charged, made up term by future monopolists. Stealing. Not distilling, not stealing “non rivalrous goods.” Just stealing. Do you agree with what he actually said?

Sticking to what was literally said and not meant by a person in a casual comment/conversation is certainly a strategy that can be used.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#282
post #138

"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge. Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economic...

Suppose you hired a consulting firm to write a report, and they delivered the report but not the internal conversations they had when developing it. You exploit a vulnerability in their phone system to get those conversations. You can argue over semantics of whether “theft” is what you did, maybe the right word is “espionage” or “spying”, but that either way we probably agree you are guilty of something? Paying for t…

But they did deliver the internal notes, just told you to not look at them. Also the analogy doesn’t make a lot of sense to me since humans (or companies paying them) own the content they produce. Based on current precedent Anthropic doesn’t have any more rights to the LLM outputs produced by your inputs than you.

Whether that violates the ToS is another matter Anthropic is of course free to sue for damages or stop doing business with you.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#283

Earlier quoted context omitted.

They definitely stole the data to make the models, but they do not say that they stole the data to make the models, but they do say that others using their outputs for unauthorized purposes is stealing. Do you see the point?

Crawling the internet and dumping it to disk is not "stealing".

If they were only copying, for example, New York Times articles and many publishers to a disk, I don't think NYT and the publishers would have sued OpenAI. But OpenAI isn't just copying things to disk. NYT reported ChatGPT (before Dec 2023, [0]) was returning near verbatim sections of NYT articles.

Is this stealing? Is it depriving NYT or publishers/writers from money via lost sales/subs? I don't know, but it certainly could be.

[0] https://www.nytimes.com/2023/12/27/business/media/new-york-t...

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#284
post #227

Earlier quoted context omitted.

This isn't a safety issue, it's LLM companies trying to be opaque and stop distillation.

By some interpretations protecting your frontier model with strong safeguards from being distilled into an open source model without safeguards is a safety issue.

There's no way to make a model "safe", (whatever that means) since you can't know what users will do with the output. It's just PR.

Closed models are also used for nefarious usage.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#285
Cool find, but can't help myself thinking that registering a domain name and submitting a paper on this to Arxiv is a bit... much. The content here could fit in a tweet or a short blog post as well. Not sure about the scientific novelty here as we're basically poking around the very top layers of someone else's software stack?

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#286
post #138

"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge. Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economic...

Suppose you hired a consulting firm to write a report, and they delivered the report but not the internal conversations they had when developing it. You exploit a vulnerability in their phone system to get those conversations. You can argue over semantics of whether “theft” is what you did, maybe the right word is “espionage” or “spying”, but that either way we probably agree you are guilty of something? Paying for t…

You paid the consulting firm for the outcome. If they sent you a bill for every piece of research they wrote down to get to the report, you bet I would want to see exactly what's inside and what I paid for.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#287
post #285

Cool find, but can't help myself thinking that registering a domain name and submitting a paper on this to Arxiv is a bit... much. The content here could fit in a tweet or a short blog post as well. Not sure about the scientific novelty here as we're basically poking around the very top layers of someone else's software stack?

Cheap investment to boost ones CV.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#288

Earlier quoted context omitted.

No, we actually don’t all know that.

I'm pretty sure essentially all HN participants understand that the frontier labs indiscriminately sucked up every bit of human output they could, IP and ethical concerns be damned. Some of that cohort may indeed be okay with it, but that doesn't change the facts.

Knowing that they trained on that data doesn't mean that you've demonstrated that they "stole" it. Certainly the courts haven't decided that in every case.

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#289

Earlier quoted context omitted.

It’s about being able to change models mid-task. For example, I want to be able to plan using Fable but implement the plan using Sonnet, and that won’t work if this is implemented.

Or even my fable credits run out mid task and need to switch back to opus >.<

Actually, that brings up a good reason they can't fix it. Fable falls back to Opus when the topic is too "unsafe". That behavior requires traces than can move between models!

Re: Stealing Reasoning Traces from Proprietary LLM APIs

#290
post #228

Earlier quoted context omitted.

This article shows that when Kimi3's chain of thought is prefilled to match Opus's, the rest of the chain of thoughts Kimi3 outputs very closely aligns with Opus's. That seems strong evidence that Kimi3 is partly a distillation of Opus. And Kimi3 is not a small model. No doubt a lot of hard work went into Kimi, but seems clear that distillation was used effectively as well. (though maybe there's another interpretatio…

Didn’t Kimi3 release a week before opus 5?

They compare it to Opus 4.8 in the article, which has been available for a few months now.
Post reply on HN