Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

301–310 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#301

Earlier quoted context omitted.

The churn is an issue. For a while there, it felt like we were getting a new number format every month or two (e.g., fp4, ternary, etc.). That level of innovation works against moving things into hardware, or at least you need to be willing to spin hardware constantly.

Even then... Do we now stop the with to asic DeepSeek and do k3 instead? Assuming such work was happening.

My personal feeling is to not move to ASICs just yet. Things are still pretty frothy right now, so I would probably wait 6-12 months. At some point the froth always calms down. At that point, commit to ASICs. That said, I’m also not totally sure exactly how much the ASIC hard codes vs having some wiggle room. The Taalas site is a bit vague as to exactly how they encode the model.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#302
post #5

I think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the mino…

The $200 you’re paying is heavily subsidized, you need to consider the real inference cost which could be closer to $10k if you’re using the maximum permitted by your plan.

May I know the source for "closer to $10k"?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#303

Earlier quoted context omitted.

I see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO. I could easily see OpenAI become…

There's a lot of value in having your product's name be synonymous with the product category. Google's quality has gone to shit, but internet searching is still "googling", both in verbiage and in the actual service people use. It's not an impenetrable moat, but OpenAI would have to stumble pretty hard to lose all of that edge.

I agree that there’s huge value in being the generic category name (see Kleenex). That said, Google and Apple are going to roll this into every phone and tablet and laptop/Chromebook. Yea, OpenAI can release a phone app (already have), but I’ll bet you dollars to donuts that the Google and Apple integrations will be superior, and even if the EU or somebody forces them to create an “AI provider neutral API” like they have for web search in browsers, most people will just roll with the default like they do for Google Search. You might be allowed to choose OpenAI, but most normies won’t.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#304

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

Your last two sentences remind me of the olden days of people building applications on top of FB and Twitter APIs (and later, Reddit), only to to have the rug pulled out from under them in one way or another

Olden days being FB and Twitter APIs? Don't forget about Apple Sherlocking independent devs. This has been happening with almost any platform for a very long time.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#305
post #276

Earlier quoted context omitted.

> US companies models constantly distill each other as Musk was forced to admit under oath > This whole narrative has just been a massive cope. So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do i…

I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow.

The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct.

Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#306

Earlier quoted context omitted.

Feel free to look, but don't bother replying if you have something negative to say, don't feel like having my day ruined by negativity. https://github.com/rdaum/mica/tree/main/crates/relation-kern... https://github.com/rdaum/mica/tree/main/crates/vm https://github.com/rdaum/pagebox/blob/main/crates/wal/src/wa...

Thanks for sharing. Looks interesting. Was all the code AI generated, or is it a mix?

that stuff in particular -- AI generated with heavy heavy prompting and up front design work and post-implementation testing

I have CUDA work here somewhere too but I have the repository private right now

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#307
post #305

Earlier quoted context omitted.

I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow.

The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct. Both sides have extremely smart people. One side has more $$$ an…

Indeed we will, my prediction is that we'll see a model from China that definitively surpasses any US model by the end of the year. China has an absolute population advantage here along with having a much better education system. And China now dominates in published AI papers.

The US enjoyed an early advantage due to excessive money being poured into AI which led to the current bubble, and access to the hardware that was needed to train these models initially.

At this point, neither of these factors actually matter that much. The naive approach of simply making models bigger has hit a wall, and now you need ingenuity in figuring out better architecture for them. Precisely because Chinese companies have had to deal with more limited resources, they put a lot more effort into researching different kinds of optimizing techniques. And of course, China is also catching up in chip making, and Huawei clusters are already competitive with Nvidia for training. So, that gap is closing as well.

The big difference is that an absolutely insane amount of money has been spent in the US, while China managed to do this on a fraction of the budget. The AI Investment Surge graph here puts things in perspective. https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#308

Earlier quoted context omitted.

And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening…

Probably a little of both. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions. Also ther…

That's my thinking as well. The whole distillation thing is a distraction from the actual innovation happening in this space. What will be interesting to see going forward is what types of new techniques people manage to come up with to over come the current architecture limits.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#309

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model m…

Apparently Google and team will build ASICs to run Gemini even though they have TPUs and it’ll be somewhat dynamic since the weights can be made modular [1]

[1] https://finance.yahoo.com/technology/ai/articles/google-plan...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#310

Earlier quoted context omitted.

Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.

Sorry, which billion-dollar corporation is Fable running "on its own"?

Fable 7 isn't a thing, 5 is the one people are currently excited about. They're talking about a hypothetical more-advanced future version.
Post reply on HN