Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

281–290 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#281

Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??

the only real benefit is privacy which 99.9% of people dont get about. Almost all serving metrics (cost, throughput, ttft) are better with large gpu clusters. Latency is usually hidden by prefill cost.

More and more people I talk to care about privacy, but not in SF

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#282

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…

That quote from Google is 2.5 years old.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#284
post #181

Earlier quoted context omitted.

How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?

If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…

Mixture-of-Expert models benefit from economies of scale, because they can process queries in parallel, and expect different queries to hit different experts at a given layer. This leads to higher utilization of GPU resources. So unless your application is already getting a lot of use, you're probably under-utilizing your hardware.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#285

Earlier quoted context omitted.

As a government contractor, using a Chinese model is a non-starter.

I don't know that it's actually prohibited. There is no Chinese telecommunications equipment allowed, no Huawei or Bytedance, but nothing prohibiting software merely being developed in China, not yet at least. Although I did just check what regions AWS bedrock support Deepseek and their govcloud regions do not, so that's a good reason not to use it. Still, on prem on a segmented network, following CMMC, probably perm…

There’s nuance and debate about the 110 level 2 controls without bringing Chinese tech in to the picture. I’d love to be a fly on the wall in that meeting lol.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#286

After using it a couple hours playing around, it is a very solid entry, and very competitive compared with the big US relaeses. I'd say it's better than GLM4.6 and I'm Kimi K2. Looking forward to v4

Did you try with 60k+ context? I found previous releases to be lacklustre which I tentatively attributed to the longer context, due to the model being trained on a lot of short context data.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#287

Earlier quoted context omitted.

I call this the "Karl Marx Fallacy." It assumes a static basket of human wants and needs over time, leading to the conclusion competition will inevitably erode all profit and lead to market collapse. It ignores the reality of humans having memetic emotions, habits, affinities, differentiated use cases & social signaling needs, and the desire to always want to do more...constantly adding more layers of abstraction in…

this name is illogical as karl marx did not commit this fallacy

Yes, he did, and it was fundamental to his entire economic philosophy: https://en.wikipedia.org/wiki/Tendency_of_the_rate_of_profit...

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#288
post #121

Earlier quoted context omitted.

Children do the same thing intuitively: parents continually complain that their children don't listen to them. But as soon as someone else tells them to "cover their nose", "chew with their mouth closed", "don't run with scissors", whatever, they listen and integrate that guidance into their behavior. What's harder to observe is all the external guidance they get that they don't integrate until their parents tell the…

Or in many cases they go over to their grandparents house and they let them run wild and all of the sudden your parents have “McDonald’s money” for their grandkids when they never had it for you.

[deleted]

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#290

Earlier quoted context omitted.

The bar is incredibly low considering what OpenAI has done as a "not for profit"

You need get a bunch of accountants to agree on what's profit first..

Agree against their best interest, mind you!
Post reply on HN