Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

321–330 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#321

Earlier quoted context omitted.

Why does that matter? They wont be making at home graphics cards anymore. Why would you do that when you can be pre-sold $40k servers for years into the future

Because Moore's law marches on. We're around 35-40 orders of magnitude from computers now to computronium. We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively…

Nothing to do with Moores Law or AGI.

The current models are simply inefficient for their capability in how they handle data.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#322
post #180

Earlier quoted context omitted.

If the Chinese model becomes better than competitors, these worries will suddenly disappear. Also, there are plenty startups and enterprises that are running fine-tuned versions of different OS models.

No… Nobody I work for will touch these models. The fear is real that they have been poisoned or have some underlying bomb. Plus y’know, they’re produced by China, so they would never make it past a review board in most mega enterprises IME.

I work at a F50 company and Deepseek is one of the model that has been approved for use. Took them a bit to get it all in place but it's certainly being used in Megacorps.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#324

Earlier quoted context omitted.

So much worse for American companies. This only means that they will be uncompetitive with similar companies that use models with realistic costs.

I can’t think of a single major US company that is big internationally that is competing on price.

> I can’t think of a single major US company that is big internationally that is competing on price.

All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#325

Earlier quoted context omitted.

As a government contractor, using a Chinese model is a non-starter.

I don't know that it's actually prohibited. There is no Chinese telecommunications equipment allowed, no Huawei or Bytedance, but nothing prohibiting software merely being developed in China, not yet at least. Although I did just check what regions AWS bedrock support Deepseek and their govcloud regions do not, so that's a good reason not to use it. Still, on prem on a segmented network, following CMMC, probably perm…

> I don't know that it's actually prohibited.

Chinese models generally aren't but DeepSeek specifically is at this point.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#326
What version is actually running on chat.deepseek.com?

It refuses to tell me when asked, only that it's been train with data up until July 2024, which would make it quite old. I turned off search and asked it for the winner of the US 2024 election, and it said it didn't know, so I guess that confirms it's not a recent model.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#327

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

I don't see what OpenAI's niche is supposed to be, other than role playing? Google seems like they'll be the AI utility company, and Anthropic seems like the go-to for the AI developer platform of the future.

Anthropic has RLed the shit out of their models to the extent that they give sub-par answers to general purpose questions. Google has great models but is institutionally incapable of building a cohesive product experience. They are literally shipping their org chart with Gemini (mediocre product), AI Overview (trash), AI Mode (outstanding but limited modality), Gemini for Google Workspace (steaming pile), Gemini on Android (meh), etc.

ChatGPT feels better to use, has the best implementation of memory, and is the best at learning your preferences for the style and detail of answers.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#328

Earlier quoted context omitted.

this name is illogical as karl marx did not commit this fallacy

Yes, he did, and it was fundamental to his entire economic philosophy: https://en.wikipedia.org/wiki/Tendency_of_the_rate_of_profit...

no, he didn't, and your link has nothing to do with your fallacy you were talking about

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#329
post #181

Earlier quoted context omitted.

How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?

If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…

> with your own hardware

Or with somebody else's.

If you don't have strict data residency requirements, and if you aren't doing this at an extremely large scale, doing it on somebody else's hardware makes much more economic sense.

If you use MoE models (al modern >70B models are MoE), GPU utilization increases with batch size. If you don't have enough requests to keep GPUs properly fed 24/7, those GPUs will end up underutilized.

Sometimes underutilization is okay, if your system needs to be airgapped for example, but that's not an economics discussion any more.

Unlike e.g. video streaming workloads, LLMs can be hosted on the other side of the world from where the user is, and the difference is barely going to be noticeable. This means you can keep GPUs fed by bringing in workloads from other timezones when your cluster would otherwise be idle. Unless you're a large, worldwide organization, that is difficult to do if you're using your own hardware.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#330
post #319

Earlier quoted context omitted.

It's like AMD open-sourcing FSR or Meta open-sourcing Llama. It's good for us, but it's nothing more than a situational and temporary alignment of self-interest with the public good. When the tables turn (they become the best instead of 4th best, or AMD develops the best upscaler, etc), the decision that aligns with self-interest will change, and people will start complaining that they've lost their moral compass.

It's not. This isn't about competition in a company sense but sanctions and wider macro issues.

It's like it in the sense that it's done because it aligns with self-interest. Even if the nature of that self-interest differs.
Post reply on HN