Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??
the only real benefit is privacy which 99.9% of people dont get about. Almost all serving metrics (cost, throughput, ttft) are better with large gpu clusters. Latency is usually hidden by prefill cost.
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
281–290 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#282How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#283Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#284Earlier quoted context omitted.
How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?
If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#285Earlier quoted context omitted.
As a government contractor, using a Chinese model is a non-starter.
I don't know that it's actually prohibited. There is no Chinese telecommunications equipment allowed, no Huawei or Bytedance, but nothing prohibiting software merely being developed in China, not yet at least. Although I did just check what regions AWS bedrock support Deepseek and their govcloud regions do not, so that's a good reason not to use it. Still, on prem on a segmented network, following CMMC, probably perm…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#286After using it a couple hours playing around, it is a very solid entry, and very competitive compared with the big US relaeses. I'd say it's better than GLM4.6 and I'm Kimi K2. Looking forward to v4
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#287Earlier quoted context omitted.
I call this the "Karl Marx Fallacy." It assumes a static basket of human wants and needs over time, leading to the conclusion competition will inevitably erode all profit and lead to market collapse. It ignores the reality of humans having memetic emotions, habits, affinities, differentiated use cases & social signaling needs, and the desire to always want to do more...constantly adding more layers of abstraction in…
this name is illogical as karl marx did not commit this fallacy
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#288Earlier quoted context omitted.
Children do the same thing intuitively: parents continually complain that their children don't listen to them. But as soon as someone else tells them to "cover their nose", "chew with their mouth closed", "don't run with scissors", whatever, they listen and integrate that guidance into their behavior. What's harder to observe is all the external guidance they get that they don't integrate until their parents tell the…
Or in many cases they go over to their grandparents house and they let them run wild and all of the sudden your parents have “McDonald’s money” for their grandkids when they never had it for you.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#289"create me a svg of a pelican riding on a bicycle"