This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge pric…
Kimi-K3 on HuggingFace
511–520 of 588 posts
Re: Kimi-K3 on HuggingFace
#512Earlier quoted context omitted.
Having worked in / adjacent several such industries, a lot of the question depends on scale. A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this. Worst-case example: Bootstrapped startup working in military. It's also the case that an open model enables many mo…
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised. I w…
The provider needs to comply with specific rules, have specific certifications, and sign specific agreements. You check the boxes, and you're good to go.
Microsoft does that better than anyone. OpenAI and Anthropic don't do that at all. Google does that rarely and poorly. AWS is not bad, but not as good as Microsoft.
Azure was always my go-to for regulated applications in the cloud. Some do require e.g. on-prem or even air gap, where even Azure is out.
Re: Kimi-K3 on HuggingFace
#513Earlier quoted context omitted.
Having worked in / adjacent several such industries, a lot of the question depends on scale. A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this. Worst-case example: Bootstrapped startup working in military. It's also the case that an open model enables many mo…
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised. I w…
Re: Kimi-K3 on HuggingFace
#514Re: Kimi-K3 on HuggingFace
#515Earlier quoted context omitted.
> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…
Around here electricity companies quote prices like yours but that is supply only while transmission, taxes, and fees are again as much on top. Is that really all inclusive?
Re: Kimi-K3 on HuggingFace
#516Earlier quoted context omitted.
I love fireworks.ai! They launched it couple of hours ago and we have it now already live on our platform for our users. Just a shame they deprecated the on-demand flux models :( Where do I get my fix for image gen now?
Fal.AI is what I use for running proprietary image models through their paces as part of my GenAI benchmark site. Highly recommended. https://fal.ai/explore
Re: Kimi-K3 on HuggingFace
#517Earlier quoted context omitted.
I'm in Arkansas and get rates fairly similar as quoted. From my last bill > KWH USAGE 2590 - $183.37 There's a base customer cost of $18 on top of that, but yeah ~$0.077/kWh taxes included.
Is that for a month or a year?
Last one was $212 for 2,146 kwh between June 8 - July 6 (28 Days)
Re: Kimi-K3 on HuggingFace
#518Earlier quoted context omitted.
There will be a huge market for local inference once it's cheap and widely available. Try to imagine output token speeds of 15,000 tok/s and a time-to-first-token of 200ms. (This has already been done for Llama 8B.) Now imagine gargantuan context windows (2M, 4M, or even bigger); keep in mind the 1M context windows were science fiction a few years ago... now imagine having this on a local model on something like a ph…
> There will be a huge market for local inference once it's cheap and widely available. I've seen public pronouncements that the RAM shortage could persist for a decade. And then if consider that the constraint on local LLMs isn't just memory size but bandwidth ... If you take something like a DGX Spark and increase its memory to 512GB that doesn't even solve the problem. Because the bandwidth of DDR5 just can't mana…
Re: Kimi-K3 on HuggingFace
#519Earlier quoted context omitted.
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised. I w…
The expectation that one BigTech company has a competent security team while the other doesn't seems entirely baseless?
Re: Kimi-K3 on HuggingFace
#520Earlier quoted context omitted.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socke…
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…