Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

241–250 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#241

Earlier quoted context omitted.

Kimi K3 is still worse than Fable and Fable was trained >4 months ago.

To say X is perfectly bad vs Y is false. People use these models for diff things. Its quite possible for the things they are used for, people do not see much of a difference. Do you hold stock in Anthropic?

> Profile created 3 days ago

> Unnecessarily aggressive

> First ever comment said "Further releases of Chinese models that demonstrate the gap is not growing substantially is a huge problem. The spending will be called into question."

Yeah I think you have an agenda

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#242

Earlier quoted context omitted.

I'd imagine it has to be immense. Most of the engineers I know are spending $200/day, let's say 20 workdays per month, $4000/mo. Whatever % of revenue it is, it's got to be close to 100% of profit.

Who is paying $200/day? My company has me on a $30/month plan (I assume they pay annually?). I use Claude constantly and only hit usage limits when I try to do 4+ projects at once. I can't imagine why anyone would need almost 7x that.

If your company is a certain size or cares about ZDR, you need to be on an Enterprise plan which is API pricing. $300/day here on average.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#243

To everyone praising Open weight models, could you answer a simple question? If Anthropic doesn't make money because of distillation attacks, how would they convince investors to invest in them, such that it makes financial sense for Anthropic to train even bigger models? Assuming it is preferable for everyone that we get better models in the future. Distillation attacks remove the financial incentive.

> f Anthropic doesn't make money because of distillation attacks, how would they convince investors to invest in them,

I dont care if that company dies. Or if OpenAI dies. I dont mind investors investing into other things either.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#244
post #5

I think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the mino…

The $200 you’re paying is heavily subsidized, you need to consider the real inference cost which could be closer to $10k if you’re using the maximum permitted by your plan.

The "real" cost is what people pay, not what the AI companies claim.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#245

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

GPUs are already ASICs.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#246

To everyone praising Open weight models, could you answer a simple question? If Anthropic doesn't make money because of distillation attacks, how would they convince investors to invest in them, such that it makes financial sense for Anthropic to train even bigger models? Assuming it is preferable for everyone that we get better models in the future. Distillation attacks remove the financial incentive.

This assumes all the open models are just a result of distilling Anthropic models. Which remains to be proven. And if they are, the point remains that Anthropic has a brittle product advantage that users and investors should be cautious about.

OpenAI says they're not just distillation. "Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that."

If you put any faith in claims made by OpenAI's "Head of Strategic Futures," then it's hard to believe these models are just the result of distillation.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#247

Earlier quoted context omitted.

How does an ASIC manage to have 60x the memory bandwidth needed to achieve that speedup?

Currently the setup is paged view in RAM shuttled to HBDRAM (VRAM) on the GPU, which in turn has to get materialized piece by piece onto cache SRAM on the GPU. Cerebras tries to get around this by keeping everything on cache SRAM as much as possible, which it burns directly to the chip wafer itself and physically places that SRAM directly next to the tiny compute unit that does the actual math. An ideal setup (not su…

I assume ROM would be even more expensive than SRAM?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#248

Earlier quoted context omitted.

Except aren’t these Chinese models having to go big too? While they might not be Mythos/Fable sized, a trillion plus parameters is hardly a local server or even desktop to mainframe style jump

I've had 2b models give a plausible Paris vacation itinerary. A tools-capable 12b and especially 30b model from 2026 is certainly capable of producing passable results. I was demonstrating the qwen 3.6 27b model I stood up last week to my wife and it gave her a passable Moroccan Chicken recipe. With tool calling (search) they're quite good.

I was on a long international flight recently with no internet and Google AI Edge Gallery installed on my phone with a 2.5GB quantized version of Gemma 4 on it. I was able to chat to it for a while and get some information about my destination which all turned out to be true and good advice.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#249
post #134

Earlier quoted context omitted.

A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand. 640K ought to be enough for anybody.

Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.

Sorry, which billion-dollar corporation is Fable running "on its own"?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#250
post #87

Earlier quoted context omitted.

ChatGPT is synonymous with non technical/work related LLMs. They're amassing a ton of user history. That history improves the product for the user because it has more context into the person. They can feed it back into model improvements and for advertising. You can see a future where a user types in "plan a vacation for me" and ChatGPT coordinates everything from there. Those sorts of users aren't going to switch be…

I see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO. I could easily see OpenAI become…

There's a lot of value in having your product's name be synonymous with the product category.

Google's quality has gone to shit, but internet searching is still "googling", both in verbiage and in the actual service people use. It's not an impenetrable moat, but OpenAI would have to stumble pretty hard to lose all of that edge.

Post reply on HN