I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models
Kimi-K3 on HuggingFace
321–330 of 588 posts
Re: Kimi-K3 on HuggingFace
#322I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models
The SemiAnalysis piece on this is long but very much worth reading: > The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole. > And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a…
https://newsletter.semianalysis.com/p/the-future-of-meta-sup...
Re: Kimi-K3 on HuggingFace
#323Did someone run censorship and political bias tests on this ? Must be interesting.
It would, indeed, be interesting to compare, given what we know about Anthropic’s censorship and political bias in their closed and more expensive models. https://x.com/dhh/status/2081435006770249831 (from the creator of Ruby on Rails). In his specific case Kimi did the task it was asked to do (translation of the article DHH wrote), which Claude refused to.
'When gypsies appropriate public spaces, you deport them. It's not hard, it's not cruel. It's the basic logic of self-protection.'
Re: Kimi-K3 on HuggingFace
#324Earlier quoted context omitted.
Yes, from 6 days ago: https://news.ycombinator.com/item?id=48986351
A question of course would be "is 15,000 tok/s Gemini better than 100 tok/s Opus 5"?
Re: Kimi-K3 on HuggingFace
#325Earlier quoted context omitted.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…
Re: Kimi-K3 on HuggingFace
#326Re: Kimi-K3 on HuggingFace
#327I feel like most hardware to run LLMs on is shaped wrong for individuals. It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards wou…
Re: Kimi-K3 on HuggingFace
#328Less than 2 hours left to the Kimi moment. It’s been more than 2 years since the DeepSeek moment that shook the world.
Re: Kimi-K3 on HuggingFace
#329Earlier quoted context omitted.
Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.
The epyc Venice SP7 socket is apparently 9324 pins https://x.com/tomshardware/status/2066846693778510331
Re: Kimi-K3 on HuggingFace
#330In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.