Live data from Hacker News

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

emergingtrajectories.com

221–230 of 349 posts

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#221

I think a big question is whether any of these labs can produce a model that is _ahead_ of Anthropic and OpenAI. A related question is how much they're dependent on the APIs of Anthropic and OpenAI to achieve their results - whether through distillation or other uses. If these models are derivative of Anthropic/OpenAI I would expect performance to be more narrow and progress to be limited.

They dont need to be ahead on performance alone.

Its value per unit of currency spent.

Financials will ultimately drive decision making.

We are already seeing that more intelligence does not correlate with more revenue, for the firm purchasing tokens.

If I was OAI/Anthropic I'd be brown and yellow in the boxers.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#222

Earlier quoted context omitted.

I see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO. I could easily see OpenAI become…

[flagged]

Rude

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#223

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

I'm convinced there is a Transmeta/transpiler solution somewhere between 9000 t/s ASIC and Nvidia GPUs. LLMs fit the model much better than what they were trying to do with general purpose in the 2000s.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#224

Earlier quoted context omitted.

[flagged]

Are you being paid by China? First you accuse me of simply speaking as a financial benefactor, then you totally shift the conversation into something that doesn't refute what I said. If you let a model spin too much, you get worse results... that's a fact, even for the best models. Nothing you've stated since actually refutes that and acting like a paid bot doesn't mean anything. I mean, if you want to felate Xi Jinp…

[flagged]

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#225
post #209

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

My gut is that, at least for now, there's a timescale mis-match problem. Hardware still takes too long, then you have to deploy it. I don't know much about "burning asics", but if the whole process of spinning up programmable GPU data centers is months, I imagine the whole ASIC cycle has some catching up to do. To be clear, by timescale mismatch I mean model quality improvement timescale vs. deployment timescale. But…

If Qwen 3 Coder were available as a PCI card I could just slot into my desktop, and the software on the system could recognize the card and still work with other cards should I decide I later want to upgrade to Qwen 5 or whatever, and at a three figure price point, I imagine it would do quite well.

The comparison I've seen elsewhere is the old console systems with separate cartridges for games... I wouldn't want to be regularly swapping them, but if they came out in a form factor that didn't require me to shell out multiple-4 digit figures in upgrades just to use the next model, it'd easily be worth it for me.

I've already got a home lab, and it's specced to last, minus the GPU. I picked up a separate system for local llm experimentation, but I'm not likely to be upgrading it again. The value add is incredibly small compared to the cost. The real benefits are data privacy and never worrying about rate limits, and there's a price point beyond which an incremental improvement to the model doesn't justify upgrading the system.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#226

Earlier quoted context omitted.

Kimi K3 is still worse than Fable and Fable was trained >4 months ago.

To say X is perfectly bad vs Y is false. People use these models for diff things. Its quite possible for the things they are used for, people do not see much of a difference. Do you hold stock in Anthropic?

ah yes, because i said something factually accurate and vaguely positive about Anthropic I must be a shareholder which would mean I either run a venture capital firm or am a current employee of Anthropic...

I wish.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#227

I keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launc…

Your last two sentences remind me of the olden days of people building applications on top of FB and Twitter APIs (and later, Reddit), only to to have the rug pulled out from under them in one way or another

Back in the day, a contract company I worked for had entire teams of people dedicated to building FB apps. When that collapsed, it came down fast and hard.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#228
post #28

Earlier quoted context omitted.

> For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the minority (correct me if I'm wrong,…

$200/mo for 20x Max is great value - but it's not sustainable and won't last forever. Microsoft gave up. Anthropic will too. I use roughly $10k/mo - at that price it's absolutely not worth it. I'm genuinely not sure where the balancing point even is.

Probably increasing learning rate have some value for these players which is enabled by subsidized usage.

And cost of acquiring learning via users feedback versus additional learning data value for a given company will dictate whether such company will give up or will continue subsidized usage, despite being unprofitable on the paper, but probably really valuable as a long term strategy for GTM and product development and increased learning rates.

Re: Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

#230

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software en…

> Does anyone think we need a Mythos level model to plan a road trip

Yes. The latest OpenAI and Anthropic models are terrible at planning roadtrips.

This is something I try to use them for frequently. They constantly get things completely wrong.

I’d say that about half of the stops they suggest fail to follow whatever filters I’ve asked for.

Post reply on HN