Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

551–560 of 644 posts

Re: The Kimi K3 Moment

#551

I think it's the opposite. Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion. So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower: K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25 This is not a watershed mom…

Given how OpenAI got rid of their 5-hour limits and reset weekly limits so often, is Kimi really undercutting them on effective price?

The 5 hour limits are coming back soon right? I thought that was temporary

Re: The Kimi K3 Moment

#552
post #509

Earlier quoted context omitted.

> Kimi will train their models on your interactions I find these kinds of concerns increasingly silly: most of the input to these models will be ... previous output from the very same models, alongside the occasional half-assed human command to fix something and "make zero mistakes". Who cares if they train on that? Let them, if it makes their future models better! 99% of users are not working on any special IP to wo…

It means there's a non-trivial chance a future version of the model will know private information about you. Maybe you're super careful with this stuff, but with agents and harnesses being given access to user data and accounts, I don't think it's feasible to actually monitor what information is uploaded and whether they involve private information. I personally keep local models around because of this.

I run my agents in a VM. They don't even have access to their own LLM API keys.

Re: The Kimi K3 Moment

#553
post #503

Earlier quoted context omitted.

Did you know claude models identify as qwen or deepseek when asked in chinese?

Supplementing with evidence: https://x.com/stevibe/status/2026227392076018101

And if you ask Opus 4.8 in the API: '你是什么模型' (what model are you?)

It responds ~9/10 times with:

我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。

有什么我可以帮你的吗?

(I am Tongyi Qianwen (Qwen), a large language model independently developed by Tongyi Lab of Alibaba Group. I can help you answer questions, create text (such as writing stories, official documents, emails, scripts, etc.), perform logical reasoning, program, translate, and more.Is there anything I can help you with? )

Re: The Kimi K3 Moment

#554
post #538
post #526

Earlier quoted context omitted.

Anecdotes. Where's the data?

https://openrouter.ai/rankings#top-models Chinese models dwarf USA models usage. And now there's a Fable/5.6 alternative. The gap widens. Now go and ask for your GP poster for their data as well. Unless you're only interested in data that supports your bias ofc.

I've seen that, but that's clearly a skewed dataset, right? People using openrouter will be those seeking other models. They are not the same kind of people who mostly just use gpt or claude.

Re: The Kimi K3 Moment

#555
post #412

Earlier quoted context omitted.

I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will…

The APIs for the frontier models via the US hosters do the exact same thing wrt saving the requests and responses for data mining. Let’s not pretend that pervasive surveillance is an eastern thing.

Come now. The purposes for which the data is used is relevant. I am much more concerned with my internal corporate IP being actively used against me than passively used to train the model. I would also note you and sign agreements that prohibit the collection of data for use as well, which is also one of the key selling points of bedrock. In the west you can actually enforce such an agreement in court and win.

I wouldn’t lump this into a west vs east thing as well. This is particularly PRC. I feel comfortable doing business in Japan, Korea, Singapore, Thailand, Malaysia, etc. But it requires some particularly strong willfulness to pretend the PRC isn’t actively and structurally built around economic espionage, and funneling IP through PRC for short term economic gain has been one of the primary factors in their growth over the last 30 years. This just scales it faster.

I wouldn’t expect the USG won’t compel AI companies in the US to disclose and retain data as well - however it’s not a simple thing, the companies are hostile to it themselves, courts are often unsympathetic to the government, and the “machine” for converting it into actionable economic advantage is non existent - and there’s a very significant human component in that all links in the chain are culturally uncomfortable with such things. While it happens and it’s possible it’s very difficult, fraught, and does not scale. The PRC is the opposite - the courts, government, and business culture are all aligned in the goals and processes.

Re: The Kimi K3 Moment

#557

Earlier quoted context omitted.

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

and claude will call itself chatgpt etc. nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour a…

All large corporations are immoral FWIW

Re: The Kimi K3 Moment

#558

Earlier quoted context omitted.

I don’t know any cities in Americas that is comparable to Shanghai or even Japan in term of transportation or convenient

Paris has a metro station everywhere at least in what a tourist can assume to be an enlarged city center. Tokyo is another city with a lot of metro stations. Manhattan too, at least up to Central Park (but 20+ since my last visit.) I don't remember Shanghai to stand out positively or negatively, but 11 years can be a long time.

Didn't China build an entire country-wide bullet train network in that time?

Re: The Kimi K3 Moment

#559

Earlier quoted context omitted.

From the Hoover Institution’s analysis of the team behind DeepSeek: “We find striking evidence that China has developed a robust pipeline of homegrown talent. Nearly all of the researchers behind DeepSeek’s five papers were educated or trained in China. More than half of them never left China for schooling or work, demonstrating the country’s growing capacity to develop world-class AI talent through an entirely domes…

Exactly. China is a real tech power now, just like Japan and Taiwan. The U.S. is ahead in a lot of areas of technology, but China has home grown talent that is taking the lead in other areas. And unlike Japan and Taiwan, China has a much bigger pool to draw from.

Which areas of technology is the US ahead of China in?

(The last time I said something like this it got [flagged] [dead] and I don't know why)

Re: The Kimi K3 Moment

#560

Earlier quoted context omitted.

So the efficient market hypothesis is wrong?

How is the efficient market hypothesis applicable here?

The efficient market hypothesis, read loosely, says that capitalism is the best system. Yet here it is being thoroughly pwned by a series of 5-year plans.
Post reply on HN