Live data from Hacker News

Kimi K2.6: Advancing open-source coding

kimi.com

351–360 of 394 posts

Re: Kimi K2.6: Advancing open-source coding

#351

Running it through opencode to their API and... it definitely seems like it's "overthinking" -- watching the thought process, it's been going for pages and pages and pages diagnosing and "thinking" things through... without doing anything. Sitting at 50k+ output tokens used now just going in thought circles, complete analysis paralysis. Might be a configuration or prompt issue. I guess I'll wait and see, but I can't…

I think this kind of overthinking is an extremely common pattern in the Chinese models. GLM's models are also very much like this.

Re: Kimi K2.6: Advancing open-source coding

#353

Earlier quoted context omitted.

> Its not anywhere close Close to what, and how are you measuring? > nobody in the USA would be spending 7 figures on infrastructure for it Au contraire, if AI had a moat it would pay for itself. They're funneling capital into infrastructure because they know it can't.

You need the infrastructure to train and run it regardless though. Kimi is great but I'm not getting the same performance from it running it on my MacBook or a 3090 as it running on a H100 or a Grace Hopper supercomputer. Pretend you did have said moat. Why wouldn't you also books infrastructure to run it on?

> Why wouldn't you also books infrastructure to run it on?

No, you wouldn't be using venture capital to overprovision your AI a hundredfold if selling AI was the end goal.

Re: Kimi K2.6: Advancing open-source coding

#355

Earlier quoted context omitted.

You really rely on ToS from Anthropic/OpenAI to know if they use your prompts or not? It's on their servers, why wouldn't they use our data?

Antropic and OpenAI are used by US businesses and government and they are audited and under contracts. If it's discovered they trained on data they shouldn't have had it will be the end of their business. On the other hand, good luck suing a Chinese company.

Not at all, Google/Meta... got caught all the time, where do you see it's the end of their business?

Re: Kimi K2.6: Advancing open-source coding

#356

Earlier quoted context omitted.

I think one of the motivations is undermining US companies. OpenAI and Anthropic are the two biggest players, and are American. Open weights models reduce the power those two big players have over the industry. If the Chinese companies tried to play by US rules and close-source their products then people would mostly use ChatGPT and Claude. So the Chinese companies don't make a ton of profit either way, but by releas…

Is Meta trying to keep the US from making as much profit with Llama? Is Google with Gemma? Microsoft with Phi? It's much simpler than some flag-waving nationalism.

Aren't Chinese open-source models actually the only ones that can compete with best proprietary/closed ones?

Re: Kimi K2.6: Advancing open-source coding

#357

There is some humor in the fact that china (of all countries) is pioneering possibly the world's most important tech via open source, while we (US) are doing the exact opposite.

It's humorous only because your expectations of China and the US are formed by Western propaganda.

Re: Kimi K2.6: Advancing open-source coding

#358

Earlier quoted context omitted.

If it’s newer and efficient then why is the api more expensive?

Price is set based on what people are willing to pay not based on actual costs.

I’d believe that if they didn’t lower limits

Re: Kimi K2.6: Advancing open-source coding

#360

Earlier quoted context omitted.

This made me chuckle. I didn't mean to dismiss ethical accountability for LLM training corpuses. It is a shame. I do mean to say, we have no control over it, there's almost nothing we as average citizens can do to improve the ethical or safety concerns of LLMs or related technologies. Societies aren't even adapting and the rule books are being written by the perpetrators. Might as well get out of it what we can while…

Wonder if stuff like this would affect it? https://github.com/p-e-w/heretic Guessing it probably would?

Neat project! I would be interested in a paper about this.

I think the tricky part with this type of technology is that, this works if the training data was not curated. What I mean is, if someone trains an LLM to simply not include key events it will not be able to reply

Not being a hater. This is neato!

Post reply on HN