Earlier quoted context omitted.
Cerebras is a programmable accelerator, but Taalas does burn the model into the gates.
Ah, right. I misremembered. Does Taalas have a meaningful advantage here then?
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
151–160 of 349 posts
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#152The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…
> the winner will be whoever burns their models to ASICs fastest. I don't understand how this works when the models are evolving so fast that your burned ASIC is outdated (or at least not top of the line) in a few weeks.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#153Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#154> Once a model is built, the biggest cost is inference Something i cant find any reliable data for, but would help for a sense of scale: How much use before its equal to training? I.e. assuming you have the training data and setup, and we only care for compute - How hours of using eg Kimi K3 / Fable, for it to equal the compute required to train it?
they used ~14.8 trillion tokens with about 2.66 million GPU hours. 14.8 * 3 = 44.4 t inference tokens.
obviously, this is back of the envelope math, but at 100t/s you would need like ~14k years. scale this to >100k GPUs and your in the hours to a couple days range.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#155> More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google. Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.
They've been doing hardware stuff as well. Both "consumer facing" bs like wearables / portables, but also more importantly chips for inference. Having a good model is one thing, being able to serve that model at good speeds and match demand is another. See Anthropic ~6months ago. Or Moonshot, they've already suspended subscriptions to their coding plans, because they can't meet demand.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#156I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#157Anthropic will get squeezed by open models for 80% of the use cases that don't require frontier capabilities and by vertical specific labs for the high value tasks that would (bio, finance, math, etc.), where smaller use case specific models will beat them on cost and speed while matching or exceeding the performance of their largest general models. Even their hail mary of being first to "AGI" will never happen becau…
None of my use cases require frontier capabilities but I still pay $200/month to a frontier lab. I value the additional time saved at more than $200/month. If I had to pay actual API rates, then I'm not sure what I would do, but it would not be an easy decision.
It became an easy decision, even the $200/month by Anthropic sounds like a bad deal.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#158Earlier quoted context omitted.
Our current incarnation of capitalism is all about monopolistic behaviors. If these LLM companies do get to the point of being able to replace employees I fully expect them to stop selling shovels and start producing the gold directly, anyone else be damned. And, frankly, this has always been the case. If a product is built on top of another service it has a limited lifespan. Either the product will be purchased or i…
And who will buy the gold? At a certain point everyone will be too poor to buy anything.
Wrong question. The real question is "Who and what will they buy with the gold?"
The golden rule is just the beginning, and no, there's nothing positive for the rest of us.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#159Earlier quoted context omitted.
Yes, the reckoning here will happen in a year or two when the (probably subsidized, maybe?) coding plans become either unavailable or much more costly. It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates. There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit,…
k3 costs will go down at least 3x within a week of the weights dropping. we'll get new quants, dspark speculators, distills and optimized kernels as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.
This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#160> More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google. Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.
The article does address what they see as the difference ("its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats") That said I think it is pretty easy to make a case that these would-be differentiators are either currently underwhelming or completely unproven (as in the case of hardware).