Live data from Hacker News

Kimi K3's Design Secret May Be in Its Thinking Traces

notes.designarena.ai

11–12 of 12 posts

Re: Kimi K3's Design Secret May Be in Its Thinking Traces

#11

Watching K3’s thought traces has been fascinating. I have a podcast transcription pipeline that uses an LLM pass to increase quality especially over proper nouns and Whisper hallucinations. Here is the trace for an episode of Deep Cover set in St Louis. Agent had no context except for what was in the transcript. This is the 4th time it had returned to this particular sentence: > 29. "emo's cup" — hmm, one more though…

why does newer Claude reject this

Re: Kimi K3's Design Secret May Be in Its Thinking Traces

#12
post #11

Watching K3’s thought traces has been fascinating. I have a podcast transcription pipeline that uses an LLM pass to increase quality especially over proper nouns and Whisper hallucinations. Here is the trace for an episode of Deep Cover set in St Louis. Agent had no context except for what was in the transcript. This is the 4th time it had returned to this particular sentence: > 29. "emo's cup" — hmm, one more though…

why does newer Claude reject this

It trips a content restriction. It’s vague, but I’m guessing because the podcasts are adjacent to/are copyrighted material.

What’s really interesting is it’s also tied to thinking level. The sweet spot for quality with low refusal is Claude 4.6 high effort, go higher in effort and and it will trip too. 4.7 is hit and miss with high and often with xhigh and above 4.8 is often. I didn’t try Fable. Sonnet wasn’t good enough quality wise.

It also seems to be related to asking it producing what it knows is copied content. I originally had it return full corrected transcripts and that hit all the time. Then I started having it return just correction lists and that works much better.

These big models recall is absolutely astounding. I built the transcript pipeline so I could search and find that episode that I kind of remember part of from a year ago. I asked fable to search the transcripts to find a particular Conan Needs a Friend episode (Katakai, as god made her!). It knew there were actually multiple episodes where that anecdote was told _before_ it had even searched. When Kimi K3 does a correction you can see in the thinking trace that it instantly recalls people and works before it reasons in to whether it’s certain that’s right.

I’m still testing, but Claude is going to get replaced for K3 when I get some time. It’s very good at this work.

Post reply on HN