Earlier quoted context omitted.
The churn is an issue. For a while there, it felt like we were getting a new number format every month or two (e.g., fp4, ternary, etc.). That level of innovation works against moving things into hardware, or at least you need to be willing to spin hardware constantly.
Even then... Do we now stop the with to asic DeepSeek and do k3 instead? Assuming such work was happening.
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
301–310 of 349 posts
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#302I think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the mino…
The $200 you’re paying is heavily subsidized, you need to consider the real inference cost which could be closer to $10k if you’re using the maximum permitted by your plan.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#303Earlier quoted context omitted.
I see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO. I could easily see OpenAI become…
There's a lot of value in having your product's name be synonymous with the product category. Google's quality has gone to shit, but internet searching is still "googling", both in verbiage and in the actual service people use. It's not an impenetrable moat, but OpenAI would have to stumble pretty hard to lose all of that edge.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#304I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…
Your last two sentences remind me of the olden days of people building applications on top of FB and Twitter APIs (and later, Reddit), only to to have the rug pulled out from under them in one way or another
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#305Earlier quoted context omitted.
> US companies models constantly distill each other as Musk was forced to admit under oath > This whole narrative has just been a massive cope. So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do i…
I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow.
Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work.
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#306Earlier quoted context omitted.
Feel free to look, but don't bother replying if you have something negative to say, don't feel like having my day ruined by negativity. https://github.com/rdaum/mica/tree/main/crates/relation-kern... https://github.com/rdaum/mica/tree/main/crates/vm https://github.com/rdaum/pagebox/blob/main/crates/wal/src/wa...
Thanks for sharing. Looks interesting. Was all the code AI generated, or is it a mix?
I have CUDA work here somewhere too but I have the repository private right now
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#307Earlier quoted context omitted.
I'm saying it's a cope to claim that the only reason Chinese models are catching up is due to distillation, while pointing out that distillation itself is in no way unique to Chinese companies. I'm sorry this was too complex of an idea for you to follow.
The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct. Both sides have extremely smart people. One side has more $$$ an…
The US enjoyed an early advantage due to excessive money being poured into AI which led to the current bubble, and access to the hardware that was needed to train these models initially.
At this point, neither of these factors actually matter that much. The naive approach of simply making models bigger has hit a wall, and now you need ingenuity in figuring out better architecture for them. Precisely because Chinese companies have had to deal with more limited resources, they put a lot more effort into researching different kinds of optimizing techniques. And of course, China is also catching up in chip making, and Huawei clusters are already competitive with Nvidia for training. So, that gap is closing as well.
The big difference is that an absolutely insane amount of money has been spent in the US, while China managed to do this on a fraction of the budget. The AI Investment Surge graph here puts things in perspective. https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa...
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#308Earlier quoted context omitted.
And given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening…
Probably a little of both. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions. Also ther…
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#309The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…
I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model m…
[1] https://finance.yahoo.com/technology/ai/articles/google-plan...
Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
#310Earlier quoted context omitted.
Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.
Sorry, which billion-dollar corporation is Fable running "on its own"?