Live data from Hacker News

Kimi-K3 on HuggingFace

huggingface.co

321–330 of 588 posts

Re: Kimi-K3 on HuggingFace

#321
post #136

I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models

Turns out sitting quietly in a room and doing math is worth more than all the money and network in the world.

Re: Kimi-K3 on HuggingFace

#322
post #260
post #136

I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models

The SemiAnalysis piece on this is long but very much worth reading: > The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole. > And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a…

he seems to be bullish on meta though and considers muse on slope than an intercept

https://newsletter.semianalysis.com/p/the-future-of-meta-sup...

Re: Kimi-K3 on HuggingFace

#323
post #12

Did someone run censorship and political bias tests on this ? Must be interesting.

It would, indeed, be interesting to compare, given what we know about Anthropic’s censorship and political bias in their closed and more expensive models. https://x.com/dhh/status/2081435006770249831 (from the creator of Ruby on Rails). In his specific case Kimi did the task it was asked to do (translation of the article DHH wrote), which Claude refused to.

we need open models because they let me dehumanize roma ppl. great take by dhh.

'When gypsies appropriate public spaces, you deport them. It's not hard, it's not cruel. It's the basic logic of self-protection.'

Re: Kimi-K3 on HuggingFace

#324

Earlier quoted context omitted.

Yes, from 6 days ago: https://news.ycombinator.com/item?id=48986351

A question of course would be "is 15,000 tok/s Gemini better than 100 tok/s Opus 5"?

Depends what you do. We have certain tasks we spend money on where Gemini 4.6 definitely is better than Opus 5.

Re: Kimi-K3 on HuggingFace

#325
post #309

Earlier quoted context omitted.

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…

They are most likely not based in the US, but converting to USD to make comparison easier.

Re: Kimi-K3 on HuggingFace

#326

Earlier quoted context omitted.

Yes, from 6 days ago: https://news.ycombinator.com/item?id=48986351

A question of course would be "is 15,000 tok/s Gemini better than 100 tok/s Opus 5"?

It's better for the billions of free users that Google serves.

Re: Kimi-K3 on HuggingFace

#327

I feel like most hardware to run LLMs on is shaped wrong for individuals. It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards wou…

RTX Spark does go up to 128gb No idea how it compares tho

Re: Kimi-K3 on HuggingFace

#329

Earlier quoted context omitted.

Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.

The epyc Venice SP7 socket is apparently 9324 pins https://x.com/tomshardware/status/2066846693778510331

We are going to need a bigger boat.

https://www.cerebras.ai/

Re: Kimi-K3 on HuggingFace

#330
post #267

In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.

There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.

https://bottlecapai.com/post/thinkingcap-qwen3-6-27b/

Post reply on HN