Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

41–50 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#41
If there's any hope for AI sovereignty and equality, we would have to either make expensive models cheap to run or make cheaper models do less work.

Making the latter happen involves either reformulating work in ways less intelligent LLMs can work better with. Or condensing intelligence into smaller models.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#42

Upper bound of AI progress - recursive self improvement. In this case AI will be responsible for building better models, making people who own datacenters the winners. Anthropic/OAI is cooked. Lower bound of AI progress - plateau. Progess is slowing, focus is on serving a meaningful peak capability at the lowest possible price. There's been news today that Google is building a Gemini chip with weights baked into sili…

Baking weights in makes a lot of sense for inference speed and power efficiency and has the added benefit of putting many end-users on the hardware refresh treadmill.

Yes, but how much would be a static Sonnet 3.5 be worth today? Its just about 2 years old. I'm not even going to ask about 3yo models like GPT4.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#43
post #29

So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?

you cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#44

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

Our current incarnation of capitalism is all about monopolistic behaviors. If these LLM companies do get to the point of being able to replace employees I fully expect them to stop selling shovels and start producing the gold directly, anyone else be damned. And, frankly, this has always been the case. If a product is built on top of another service it has a limited lifespan. Either the product will be purchased or i…

I mean Claude IS replacing people. Even in Australia - Claude + postgres mcp is replacing data/bi people in even small and medium businesses.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#45

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

Then run your own fine tuned models for your AI startup.

Doesn't that mean you bet against the bitter lesson?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#46

To everyone praising Open weight models, could you answer a simple question? If Anthropic doesn't make money because of distillation attacks, how would they convince investors to invest in them, such that it makes financial sense for Anthropic to train even bigger models? Assuming it is preferable for everyone that we get better models in the future. Distillation attacks remove the financial incentive.

This assumes all the open models are just a result of distilling Anthropic models. Which remains to be proven. And if they are, the point remains that Anthropic has a brittle product advantage that users and investors should be cautious about.

> investors should be cautious about

if they are cautious, what would make them invest in newer bigger models without the expected return? generosity?

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#47
post #26
post #9

Nope, I'll still buy Claude because the overall XP is better than Kimi and Qwen who literally copied basic harnesses to make kimi-cli and qwen-cli, respectively. Also, you can tell if a model is genuinely powerful and well-thought-out vs a model that acts like it. It's like Apple vs Xiaomi/Huawei. Sure, you can get a Huawei with bells and whistles, but most people learnt the hard way that those companies just copy th…

> so might as well get the real deal. Almost the same thing for twice the price, just for the pleasure of saying that you believe Apple was first?

To be fair, the flagship Xiaomi, Vivo, Oppo etc are comparable in price to other flagship devices. You very much pay a premium for the large camera sensors they put in these devices, amongst other things.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#48
The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest.

The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release.

Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software engineering (if not already). Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe?

A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand.

Furthermore, if you're an enterprise the risk of data exfiltration and feeding data to a potential competitor like OpenAI or Anthropic is greatly reduced if you could shift to on-prem ASIC deployments. A handful of chips could cover a wide variety of use cases and cover them more securely. There are a lot of corporate use-cases for LLMs that are not frontier math research or coding.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#49

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

> I would suggest to people using LLMs: you should be cautious about giving these companies data or relying on them. If you're building an AI startup, there's a very good chance they could decide to directly compete with you if your idea has traction. You're also at their mercy for API pricing etc.

LLM generated code is not copyrightable, so even if they do "steal" it - I don't think there's legal grounds to do anything about it.

It can't be stolen. You don't really own it.

You can try to lock it in a safe and hope no one ever gets a hold of it. You can lie and say you didn't use an LLM, but Anthropic and OpenAI et al probably have logs to disprove that.

But, if push comes to shove, you don't own it...

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#50
It's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything'

The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words)

Post reply on HN