Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

151–160 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#151
post #122

Earlier quoted context omitted.

Cerebras is a programmable accelerator, but Taalas does burn the model into the gates.

Ah, right. I misremembered. Does Taalas have a meaningful advantage here then?

Taalas claims significant cost, speed, and power differential with running on Nvidia (and similar) hardware. Their whole thesis is that with making the hardware cost less, run faster, and run with less power mixed with their claimed fast production cycle for new weights, companies would plan on purchasing and then relegating the other chips to different workloads at the next update cycle. And the chips have facilities for fine-tunes that can by dynamically loaded. The info I've seen has been very light on details, but if they can even get halfway where they are planning, it'll be a massive shift.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#152

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

> the winner will be whoever burns their models to ASICs fastest. I don't understand how this works when the models are evolving so fast that your burned ASIC is outdated (or at least not top of the line) in a few weeks.

for many purposes they're good enough now. If I had an opus 4.8 class model on a box next to me that could produce tokens at rates like 5000/s, i don't know if i'd need a new one for a long time. I think we might be underestimating how powerful very very fast LLMs could be, since they could iterate on tons of small variations on tasks. paired with deterministic guardrails that gate "doneness", you could loop on tasks for a long time having the agent try different strategies until the goal was reached, in ways that are just impractical now (and very expensive)

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#153
post #141

Earlier quoted context omitted.

Ah, right. I misremembered. Does Taalas have a meaningful advantage here then?

6x. Taalas has Llama3.1 8B running at 18000 tok/s. Cerebras advertised that model at 3000 tok/s.

8B is still three orders smaller than current frontier though

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#154

> Once a model is built, the biggest cost is inference Something i cant find any reliable data for, but would help for a sense of scale: How much use before its equal to training? I.e. assuming you have the training data and setup, and we only care for compute - How hours of using eg Kimi K3 / Fable, for it to equal the compute required to train it?

if you assume that training requires about 3x the compute of inference (one forward pass, one backward pass, parameter updates), and we take DeepSeek-V3 since their numbers are public.

they used ~14.8 trillion tokens with about 2.66 million GPU hours. 14.8 * 3 = 44.4 t inference tokens.

obviously, this is back of the envelope math, but at 100t/s you would need like ~14k years. scale this to >100k GPUs and your in the hours to a couple days range.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#155
post #2

> More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google. Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.

They've been doing hardware stuff as well. Both "consumer facing" bs like wearables / portables, but also more importantly chips for inference. Having a good model is one thing, being able to serve that model at good speeds and match demand is another. See Anthropic ~6months ago. Or Moonshot, they've already suspended subscriptions to their coding plans, because they can't meet demand.

Didn't the Apple lawsuit throw a monkey wrench into their hardware plans?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#156

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

[dead]

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#157
post #3

Anthropic will get squeezed by open models for 80% of the use cases that don't require frontier capabilities and by vertical specific labs for the high value tasks that would (bio, finance, math, etc.), where smaller use case specific models will beat them on cost and speed while matching or exceeding the performance of their largest general models. Even their hail mary of being first to "AGI" will never happen becau…

None of my use cases require frontier capabilities but I still pay $200/month to a frontier lab. I value the additional time saved at more than $200/month. If I had to pay actual API rates, then I'm not sure what I would do, but it would not be an easy decision.

I thought similarly until I decided to try DeepSeek.

It became an easy decision, even the $200/month by Anthropic sounds like a bad deal.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#158

Earlier quoted context omitted.

Our current incarnation of capitalism is all about monopolistic behaviors. If these LLM companies do get to the point of being able to replace employees I fully expect them to stop selling shovels and start producing the gold directly, anyone else be damned. And, frankly, this has always been the case. If a product is built on top of another service it has a limited lifespan. Either the product will be purchased or i…

And who will buy the gold? At a certain point everyone will be too poor to buy anything.

> And who will buy the gold?

Wrong question. The real question is "Who and what will they buy with the gold?"

The golden rule is just the beginning, and no, there's nothing positive for the rest of us.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#159
post #111

Earlier quoted context omitted.

Yes, the reckoning here will happen in a year or two when the (probably subsidized, maybe?) coding plans become either unavailable or much more costly. It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates. There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit,…

k3 costs will go down at least 3x within a week of the weights dropping. we'll get new quants, dspark speculators, distills and optimized kernels as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.

I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3.

This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#160
post #2

> More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google. Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.

The article does address what they see as the difference ("its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats") That said I think it is pretty easy to make a case that these would-be differentiators are either currently underwhelming or completely unproven (as in the case of hardware).

Right. Arguing with the original article, not you, my counterpoint would be, “Sure, if they’re able to do that. But so could any of the other model-only competitors.” The best you can say today is that they have announced an intention to do those things, but in no way have they established themselves as being successful, yet. And the model-only Chinese providers have access to lots of e.g. wearable and consumer tech.
Post reply on HN