Live data from Hacker News

Kimi K3-256k

kimi.com

81–90 of 172 posts

Re: Kimi K3-256k

#82
post #2

[flagged]

This exact same comment is the top comment on the reddit thread for this news https://www.reddit.com/r/kimi/s/BFa1TR9vNg

I searched one other comment of theirs in their history and see the same thing. On mobile so I can’t easily tell if this account is the bot account or if the ones on Reddit are.

HN: https://news.ycombinator.com/item?id=47934437

Reddit: https://www.reddit.com/r/worldnews/comments/1sxzzop/comment/...

Re: Kimi K3-256k

#83
post #77

Earlier quoted context omitted.

That is true, but I feel like sometimes if the conversation contains useful rational it can help to keep it. I think sometimes it is a judgement call, I will sometimes compress the context first. If I feel like the model and I explored a lot of options I won't want to keep context as it might be confusing. I think the more you use it the better judge you are of whether you should purge, compress, or just keep the con…

Models work best when they have short instructions and no noise. "Conversation [may] contain" also means "conversation has a lot of noise". It degrades performance and increases cost. I use handoff skill to ask model to write a prompt for itself.

The applicability of your advice is very model dependent. Some like Claude have very good long context performance, whereas others they fall off much quicker past some threshold.

Re: Kimi K3-256k

#84
post #11

Since Claude is the first time for me really, really out (TIL against my wished about https://status.claude.com/ ), I am now interested enough to see what else works. But ... when I click pricing, I see "Join a waitlist". Wtf? Are they really that good, so were totally surprised and overwhelmed by the requests, is this a marketing stunt, or do they just don't have the hardware being in china?

Not at all a marketing gimmick, the demand is simply that high.

I was able to press "Join waitlist" and then within 48 hours got accepted. The limits aren't very high, no where near the endless subsided+resets given on ChatGPT/Claude. I recommend ChatGPT for good value output!

Others here mentioned "the providers could be quantizing it!" but some of the providers on OpenRouter have partnered with Moonshoot and OpenRouter shows the int when you expand on the provider.

It should say "mxfp4" but providers like Baseten report FP8.

Re: Kimi K3-256k

#85

Why are Anthropic and OpenAI even allowing their coding harness apps to be plugged into different model providers…? I’m surprised they haven’t figured out a way to clamp down on that by now.

It honestly should be easier, its a by product of everyone using the same API standard.

Codex and Claude require editing a .json file, but most other harnesses have direct connections via a /provider or /login command.

Re: Kimi K3-256k

#86
Not relevant to this link but I was thinking about the allegations of Chinese AI companies distilling from the big frontier American ones. And I came to the conclusion: I don’t care.

Who cares? China has always copied and then copied the means of production and then out produced. See also Tesla and now all the Chinese cars eating their lunch.

As long as I get really solid AI models for cheap that do what I need I don’t care if they’re Chinese or otherwise.

I’ll still never use Grok from SpaceX AI cuz eww no, I have principles. ;-)

Re: Kimi K3-256k

#87

LLMs is quickly became commodities. US AI labs like OpenAI is losing their moat. Hyperscalers and data center owners who can sell cheap token will win

I have so much gratitude to the frontier companies who did all the extremely complicated research and development, model by model. It already feels difficult to remember how much capital it really took. Thank you for getting us to this point.

Re: Kimi K3-256k

#89
post #77

Earlier quoted context omitted.

Models work best when they have short instructions and no noise. "Conversation [may] contain" also means "conversation has a lot of noise". It degrades performance and increases cost. I use handoff skill to ask model to write a prompt for itself.

The applicability of your advice is very model dependent. Some like Claude have very good long context performance, whereas others they fall off much quicker past some threshold.

[deleted]

Re: Kimi K3-256k

#90
post #50

Hopefully this helps reduce some of the pressure on their infrastructure. Their models have all become super dumb recently and their support are not addressing it. I have a hunch they’ve been serving a significant percentage of requests with quantised models.

Ah come on, HN is not the place for these types of unfounded conspiracy theories that keep popping up on Reddit.
Post reply on HN