Earlier quoted context omitted.
The word "Germany" is "Germany" everwhere and always ends with y. The country Germany has different names though.
Germans prefer Deutschland, no?
Kimi K3-256k
81–90 of 172 posts
Re: Kimi K3-256k
#82[flagged]
This exact same comment is the top comment on the reddit thread for this news https://www.reddit.com/r/kimi/s/BFa1TR9vNg
HN: https://news.ycombinator.com/item?id=47934437
Reddit: https://www.reddit.com/r/worldnews/comments/1sxzzop/comment/...
Re: Kimi K3-256k
#83Earlier quoted context omitted.
That is true, but I feel like sometimes if the conversation contains useful rational it can help to keep it. I think sometimes it is a judgement call, I will sometimes compress the context first. If I feel like the model and I explored a lot of options I won't want to keep context as it might be confusing. I think the more you use it the better judge you are of whether you should purge, compress, or just keep the con…
Models work best when they have short instructions and no noise. "Conversation [may] contain" also means "conversation has a lot of noise". It degrades performance and increases cost. I use handoff skill to ask model to write a prompt for itself.
Re: Kimi K3-256k
#84Since Claude is the first time for me really, really out (TIL against my wished about https://status.claude.com/ ), I am now interested enough to see what else works. But ... when I click pricing, I see "Join a waitlist". Wtf? Are they really that good, so were totally surprised and overwhelmed by the requests, is this a marketing stunt, or do they just don't have the hardware being in china?
I was able to press "Join waitlist" and then within 48 hours got accepted. The limits aren't very high, no where near the endless subsided+resets given on ChatGPT/Claude. I recommend ChatGPT for good value output!
Others here mentioned "the providers could be quantizing it!" but some of the providers on OpenRouter have partnered with Moonshoot and OpenRouter shows the int when you expand on the provider.
It should say "mxfp4" but providers like Baseten report FP8.
Re: Kimi K3-256k
#85Why are Anthropic and OpenAI even allowing their coding harness apps to be plugged into different model providers…? I’m surprised they haven’t figured out a way to clamp down on that by now.
Codex and Claude require editing a .json file, but most other harnesses have direct connections via a /provider or /login command.
Re: Kimi K3-256k
#86Who cares? China has always copied and then copied the means of production and then out produced. See also Tesla and now all the Chinese cars eating their lunch.
As long as I get really solid AI models for cheap that do what I need I don’t care if they’re Chinese or otherwise.
I’ll still never use Grok from SpaceX AI cuz eww no, I have principles. ;-)
Re: Kimi K3-256k
#87LLMs is quickly became commodities. US AI labs like OpenAI is losing their moat. Hyperscalers and data center owners who can sell cheap token will win
Re: Kimi K3-256k
#88Re: Kimi K3-256k
#89Earlier quoted context omitted.
Models work best when they have short instructions and no noise. "Conversation [may] contain" also means "conversation has a lot of noise". It degrades performance and increases cost. I use handoff skill to ask model to write a prompt for itself.
The applicability of your advice is very model dependent. Some like Claude have very good long context performance, whereas others they fall off much quicker past some threshold.
Re: Kimi K3-256k
#90Hopefully this helps reduce some of the pressure on their infrastructure. Their models have all become super dumb recently and their support are not addressing it. I have a hunch they’ve been serving a significant percentage of requests with quantised models.