Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

321–330 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#321

Earlier quoted context omitted.

> whoever burns their models to ASICs fastest. There is already custom hardware see cerebras. GPUs have a lot of slack there is at least one lab that had a (small 8b) model generate almost 3000 tokens per second on a MI300X for a talk, instead of the typical software stack that did maybe 100ish tokens per second. High bandwidth flash storage is in the works, i.e hard drives with TBs of storage and over 1 TB per secon…

> Meaning that in a couple of years you may There is no "may" here. You will see this. It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame th…

Oh yeah the may is on the 2 year time horizon. It could be 3 or 4. Or next year.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#322

Earlier quoted context omitted.

Sunk cost fallacy. So much money has been invested in Anthropic and OpenAI at this point that to declare it a loss and walk away could potentially destroy a lot of VC firms, and a non-trivial chunk of the US Economy.

Sunk cost fallacy is not relevant to future investment.

> Sunk cost fallacy is not relevant to future investment

Outside of future choices, what in the world is sunk cost fallacy in fact relevant to? Do you understand words?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#323

Earlier quoted context omitted.

I admit I would like a faster model - but even though I have faster models available I still go to Fable or GPT-5.6 90% of the time. So there is a gap between a potential preference and a revealed preference. Custom AI for things like facial recognition in cameras has existed for decades, before LLMs were a thing. I don't see that getting replaced. And on-device conversational intelligence might go that route as well…

> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time What about all the things you don't currently use an LLM for? If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little t…

or, hear me out, advertisers can do real-time advertising based on hyper-now context

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#324
I have a problem with the cost per task metrics of Artificial Analysis. We don’t know how they calculate it exactly. But recently, cost per task has become the most discussed topic. The logic is basically: if model A achieves 55% on benchmark X and model B 60%, but the cost per task of A is 50% cheaper, people would choose A instead of B.

But that implies that all output of the less intelligent model A is usable, perhaps only a bit worse than the output of B. But what if the output of A is unusable, or it can only deliver usable results in 1 out of 5 tries? In such cases, the user will have to rerun the task and it will very quickly double or triple the cost and makes the old average number misleading! I would argue the retry and flaky cost will be many times bigger than the average token cost and that is the true cost the users have to bear.

AI-Benchy [0] (admittedly a one man benchmark) shows a much different figure than the numbers of Artificial Analysis. Opus 4.8 cost per task according to AA is $1.80 and Kimi K3 is $0.94$. According to AI Benchy, however, the *cost per successful task* of Opus 4.8 is 10.7 cents vs 19.4 cents of K3. The number of correct tests and pass rate of Opus 4.8 is also higher than Kimi K3.

So on a cost-per-usable-result basis, Kimi K3 is actually pricier than Opus 4.8 — the opposite of what AA’s headline number suggests.

Thus, I don’t know if I can believe the numbers of AA or we need to track the cost ourselves.

[0] https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#325

Earlier quoted context omitted.

I admit I would like a faster model - but even though I have faster models available I still go to Fable or GPT-5.6 90% of the time. So there is a gap between a potential preference and a revealed preference. Custom AI for things like facial recognition in cameras has existed for decades, before LLMs were a thing. I don't see that getting replaced. And on-device conversational intelligence might go that route as well…

> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time What about all the things you don't currently use an LLM for? If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little t…

Exactly, I think a lot of people aren't thinking about it this way, they're imagining a faster version of ChatGPT. In reality if it was a frontier model running a these speeds it would change so much about how we interact with computers. It would be custom hyper specific software on demand.

Looking for a lamp in a specific style?

"Make a VR application set inside my apartment (based on all the photos of my apartment from my photos directory) with all lamps under $100 that fit Scandinavian interiors and could be delivered to my house before Friday. Place the lamp on the dining room table, allow us to: cycle through lamps, change time of day, and interact with all the lamps and furniture".

Five seconds later and bam you have this new piece of software that you'll use once and then dispose of.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#326
post #268

Earlier quoted context omitted.

I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3. This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.

There's been a price war going on openrouter between providers of GLM 5.2. NovitaAI, DeepInfra, and StreamLake keeps underbidding each other in waves. Yesterday evening both input and output $/M was ~$0.3. Output was especially cheap.

I had the opposite experience. I bought the $100 monthly sub from Neuralwatt last month because it was the only economical provider for GLM 5.2. They raised their rates halfway through, and it simultaneously became too slow to use.

I just looked at DeepInfra -- I've got an account there already etc -- and it's at FP4 quant. How much that effects the quality of inference for GLM, I can't say. I could see using it as a backup when other things run out but don't think I'd trust it.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#327
post #317

Earlier quoted context omitted.

The "640k should be enough for anyone" quote (even if Gates didn't exactly say it) is making the point is that _right now_ we have no idea about what our future needs and capabilities will be, we can't imagine what "should be enough" will be 640k was enough ... in 1981 ... almost fifty years later is 50,000 lower than a standard off the shelf PC now

If that’s the case I still don’t get it. 640k was enough in 1981, same as how 150 (or 9k) tok/sec would suffice for 2026 The comment seems like the nerd equivalent of 6 7

The original quote was supposedly Bill Gates saying that "640kb of memory should be enough for anyone".

The quote became famous because it's a failure to imagine that people would find new uses for computer memory if it became plentiful and cheap. In the quote Gates is not expecting people to come up with more demanding applications for computer memory (RAM) than the ones which were available in 1981.

Computer memory did in fact become plentiful and cheap after this, and computer programs became more complex and memory-hungry and today we wouldn't consider a 640kb an acceptable amount of ram for even the lowest-end device.

OP is repeating the quote about memory to indicate that they think that modern LLMs will be a similar resource. We should not expect that demand for LLM inference will stay flat, and that once everyone has cheap access to Fable level, we won't find new more demanding uses for it.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#329

Earlier quoted context omitted.

The tokens per second performance numbers coming from Cerebrus/Talas are several orders of magnitude higher than models running on GPUs, which is such a huge step change that it will enable many more uses of LLMs that are impractical otherwise. I.e. think about gamers and burning in an LLM chip on a game console like a future Play Station - it doesn't matter if its a frontier LLM if it allows them to talk to in game…

I admit I would like a faster model - but even though I have faster models available I still go to Fable or GPT-5.6 90% of the time. So there is a gap between a potential preference and a revealed preference. Custom AI for things like facial recognition in cameras has existed for decades, before LLMs were a thing. I don't see that getting replaced. And on-device conversational intelligence might go that route as well…

You will be very happy once OpenAI serves GPT-5.6-Sol on Cerebus at 750 token/s.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#330

Earlier quoted context omitted.

Thanks for sharing. Looks interesting. Was all the code AI generated, or is it a mix?

that stuff in particular -- AI generated with heavy heavy prompting and up front design work and post-implementation testing I have CUDA work here somewhere too but I have the repository private right now

Very cool. Thanks for sharing.
Post reply on HN