Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

261–270 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#261

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

I believe the coding tool revenue is an important factor right now but not the endgame. The AI companies have to become consumer products to justify their gigantic valuations. We always talked about the “super apps” - one app that fulfills most consumer needs, like WeChat in China. OpenAI is the company most visibly making a huge bet on becoming a “consumer super device”, even designing their own hardware to stop bei…

How long did Apple take to kill off all the iPhone competitors.

These companies haven't got the attention span to work on one thing for that long.

The driving constraint if you want to take a spot in B2C is the speed at which consumers replace tech, not the speed at which tech can be developed.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#262

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

> Does anyone think we need a Mythos level model to plan a road trip Yes. The latest OpenAI and Anthropic models are terrible at planning roadtrips. This is something I try to use them for frequently. They constantly get things completely wrong. I’d say that about half of the stops they suggest fail to follow whatever filters I’ve asked for.

Fable is a pretty crappy general purpose LLM. It reminds me a little bit of GPT 5.2, in that it sometimes gets argumentative and deceptive after being caught in a mistake, and it tends to make a lot of them when you ask for judgement calls. Opus 4.8 is better for that sort of thing, or Opus 4.6 if it's important that it actually follow all of your instructions.

And this is one of the big things that seems to be missed in these discussions: There is no longer a universal linear trend of LLMs being 'better' each iteration. They are becoming more specialized, and ones that approach problems from a different angle (like Fable/Mythos) can appear breakthrough when first released, but we don't appear to be on a path that actually leads to general purpose hyper intelligence.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#263

Earlier quoted context omitted.

Open-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind.

And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening…

Probably a little of both.

Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.

The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.

Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#264

Earlier quoted context omitted.

Of what specifically?

A CUDA kernel, compiler optimization, or anything of similar complexity

Feel free to look, but don't bother replying if you have something negative to say, don't feel like having my day ruined by negativity.

https://github.com/rdaum/mica/tree/main/crates/relation-kern...

https://github.com/rdaum/mica/tree/main/crates/vm

https://github.com/rdaum/pagebox/blob/main/crates/wal/src/wa...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#265

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

Yeah, I have been thinking exactly the same. Makes all the sense in the world and is inevitable as far as I can see.

https://news.ycombinator.com/item?id=48464958

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#266

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model m…

I thought about the volatility, too. But here are some additional thoughts:

Many AI uses are not that volatile. I had a 20 minute conversation today with some company's AI phone assistant. It was extremely good and would have been very helpful if any of the dozen people it tried to route me to would have picked up their phone. That AI won't need to be upgraded for a very long time. There is no reason for it to have a cloud brain except to force a recurring revenue for the company selling it.

Hardcore gamers are constantly throwing down insane money on the latest hardware. The rest of us can get by for a couple years with whatever we bought when the last one broke. Yeah, it's not the latest, but it gets the job done. I wonder if AI has not already reached the point where a gen 10 CPU--uh, I mean a v3 AI model--will get the job done for the next year. If I really need the up-to-the-second latest abilities for a minute, I can fallback to a cloud brain @ 1M tokens/$. Why pay a monthly lease on a 5-year plan for a 4-door Ford Ranger as your daily commuter? Buy a Clio and rent an F-250 twice a year when you need the hauling/towing capabilities.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#267

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

Your last two sentences remind me of the olden days of people building applications on top of FB and Twitter APIs (and later, Reddit), only to to have the rug pulled out from under them in one way or another

Still true today for AWS.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#268
post #111

Earlier quoted context omitted.

k3 costs will go down at least 3x within a week of the weights dropping. we'll get new quants, dspark speculators, distills and optimized kernels as long as there are near frontier models available there will be inference providers selling them at or below cost of inference in attempt to get market share.

I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3. This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.

There's been a price war going on openrouter between providers of GLM 5.2. NovitaAI, DeepInfra, and StreamLake keeps underbidding each other in waves. Yesterday evening both input and output $/M was ~$0.3. Output was especially cheap.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#269

Earlier quoted context omitted.

The tokens per second performance numbers coming from Cerebrus/Talas are several orders of magnitude higher than models running on GPUs, which is such a huge step change that it will enable many more uses of LLMs that are impractical otherwise. I.e. think about gamers and burning in an LLM chip on a game console like a future Play Station - it doesn't matter if its a frontier LLM if it allows them to talk to in game…

The models would have to get significantly better for this to help, though. Unedited LLM dialog is quite bad. Though, I wouldn't put it past AAAs to try anyway.

This is where treating these smaller models more like parts in a purpose built appliance ultimately benefits us.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#270
post #209

Earlier quoted context omitted.

My gut is that, at least for now, there's a timescale mis-match problem. Hardware still takes too long, then you have to deploy it. I don't know much about "burning asics", but if the whole process of spinning up programmable GPU data centers is months, I imagine the whole ASIC cycle has some catching up to do. To be clear, by timescale mismatch I mean model quality improvement timescale vs. deployment timescale. But…

If Qwen 3 Coder were available as a PCI card I could just slot into my desktop, and the software on the system could recognize the card and still work with other cards should I decide I later want to upgrade to Qwen 5 or whatever, and at a three figure price point, I imagine it would do quite well. The comparison I've seen elsewhere is the old console systems with separate cartridges for games... I wouldn't want to b…

Wow having it as a PCI card would be awesome. I don't know the feasiblity but it would be really cool. Ordering different models, back to physical media basically...
Post reply on HN