Infrastructure owners with access to the cheapest energy will be the long run winners in AI.
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
101–110 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#102How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
Google would love a cheap hq model on its surfaces. That just helps Google.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#103Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.
How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#104I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Germany was right in some ways and wrong in others for the soviet unions strength. USSR failed to conquer Finland because of the military purges. German intelligence vastly under-estimated the amount of tanks and general preparedness of the Soviet army (Hitler was shocked the soviets had 40k tanks already). Lend Lease act really sent an astronomical amount of goods to the USSR which allowed them to fully commit to the war and really focus on increasing their weapon production, the numbers on the amount of tractors, food, trains, ammunition, etc. that the US sent to the USSR is staggering.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#105Earlier quoted context omitted.
I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…
I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#106Earlier quoted context omitted.
and can be faster if you can get an MOE model of that
"Mixture-of-experts", AKA "running several small models and activating only a few at a time". Thanks for introducing me to that concept. Fascinating. (commentary: things are really moving too fast for the layperson to keep up)
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#107To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.
Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.
This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it.
It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may become unsustainable before the last players fold.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#108How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
Pure models clearly aren’t the monetizing strategy, use of them on existing monetized surfaces are the core value. Google would love a cheap hq model on its surfaces. That just helps Google.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#109How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#110Earlier quoted context omitted.
If you thought DeepSeek "just wasn't good," there's a good chance you were running it wrong. For instance, a lot of people thought they were running "DeepSeek" when they were really running some random distillation on ollama.
WDYM? Isn't https://chat.deepseek.com/ the real DeepSeek?
I ran the 1.58-bit Unsloth quant locally at the time it came out, and even at such low precision, it was super rare for it to get something wrong that o1 and GPT4 got right. I have never actually used a hosted version of the full DS.