Earlier quoted context omitted.
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…
They are most likely not based in the US, but converting to USD to make comparison easier.
Kimi-K3 on HuggingFace
391–400 of 588 posts
Re: Kimi-K3 on HuggingFace
#392from the license: If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
good find! This sounds a bit like what Meta was doing with the earlier Llama models? There is also this paragraph in their licence that is smart marketing-wise: > 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly re…
Re: Kimi-K3 on HuggingFace
#393Earlier quoted context omitted.
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap…
They are most likely not based in the US, but converting to USD to make comparison easier.
Re: Kimi-K3 on HuggingFace
#394Earlier quoted context omitted.
Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open
That's a summarized and filtered view of the actual reasoning. OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.
Re: Kimi-K3 on HuggingFace
#395Earlier quoted context omitted.
As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.
Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection…
Re: Kimi-K3 on HuggingFace
#396from the license: If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
good find! This sounds a bit like what Meta was doing with the earlier Llama models? There is also this paragraph in their licence that is smart marketing-wise: > 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly re…
Re: Kimi-K3 on HuggingFace
#397Re: Kimi-K3 on HuggingFace
#398Earlier quoted context omitted.
It all depends if you count the fixed cost of training or not. And the cost of the hardware.
Or even the basis of the cost of hardware. There are lease deals, capacity traded for equity, various programs by Nvidia, there's absolutely massive depreciation, etc.
Everyone keeps thinking those A100s only have 6 more months of life, and yet they're still going for more than they did per hour in 2024.
Show me evidence that A100 prices have collapsed, and maybe GPU depreciation will be relevant to the market.
Re: Kimi-K3 on HuggingFace
#399Earlier quoted context omitted.
I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens. The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it. Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and special…
I don't think the goal of this data is necessarily model improvement. I think it's marketing, advertising, and product refinement. Ex: all the things Google wants your search data for. It's somewhat silly to think the value of that data has changed much. Advertisers want to know what's popular and getting clicks and attention. Competitors want to know what features are getting used in their markets. In the simplest c…
Re: Kimi-K3 on HuggingFace
#400how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?