Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

221–230 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#221

Earlier quoted context omitted.

As much I agree with your sentiment, but I doubt the intention is singular.

The bar is incredibly low considering what OpenAI has done as a "not for profit"

You need get a bunch of accountants to agree on what's profit first..

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#222
post #211
post #13

Worth noting this is not only good on benchmarks, but significantly more efficient at inference https://x.com/_thomasip/status/1995489087386771851

Do we know why?

Sparse Attention, it's the highlight of this model as per the paper

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#223
post #179

Earlier quoted context omitted.

[flagged]

Isn’t every country by definition a “local monopoly on force”? Sweden and Norway have their own militaries and police forces and neither would take kindly to an invasion from the other. By your definition this makes them adversaries or enemies.

Exactly. I am Norwegian myself, and I don’t even know how many wars we have had with Sweden and Denmark.

If you are getting at the fact that it is sometimes beneficial for adversaries to collaborate (e.g., the prisoner dilemma) then I agree. And indeed, both Norway and Sweden would be completely lost if they declared war on the other tomorrow. But it doesn’t change the fundamental nature of the relationship.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#224

Earlier quoted context omitted.

Why does that matter? They wont be making at home graphics cards anymore. Why would you do that when you can be pre-sold $40k servers for years into the future

Because Moore's law marches on. We're around 35-40 orders of magnitude from computers now to computronium. We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively…

> If we do get to AGI (2029 according to Kurzweil)

if you base your life on Kurzweil's hard predictions you're going to have a bad time

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#225
post #180

Earlier quoted context omitted.

No… Nobody I work for will touch these models. The fear is real that they have been poisoned or have some underlying bomb. Plus y’know, they’re produced by China, so they would never make it past a review board in most mega enterprises IME.

People say that, but everyone, including enterprises, are constantly buying Chinese tech one way or another because of cost/quality ratio. There’s a tipping point in any excel file where risks don’t make sense, if the cost is 20x for the same quality. Of course you’ll always have exceptions (government, military and etc.), but for private, winner will take it all.

What Chinese built infrastructure tech where information can be exfiltrated or cause any real damage are American companies buying? Chinese communication tech is for the most part not allowed in any American technology.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#226

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> Infrastructure owners with access to the cheapest energy will be the long run winners in AI.

For a sufficiently low cost to orbit that may well be found in space, giving Musk a rather large lead. By his posts he's currently obsessed with building AI satellite factories on the moon, the better to climb the Kardashev scale.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#227
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

[flagged]

There used to be memes „open source is communism”, vide https://souravroy.com/2010/01/01/is-open-source-pro-communis...

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#228

Earlier quoted context omitted.

Yeah that’s not how Big Enterprise works… And most startups are just doing prompt engineering that will never go anywhere. The big companies will just throw a couple of developers at the feature and add it to their existing business.

Big enterprise with mostly private companies as their clients? Lol, yeah, that’s how they work from my personal experience. The reality is, if it’s not a tech-first enterprise and already outsource part of tech to a shop outside of NA (which is almost majority at this point), they will do absolutely everything to cut the costs.

I spent three years working in consulting mostly in public sector and education and the last two working with startups to mid size commercial interest and a couple of financial institutions.

Before that I spent 6 years working between 3 companies in health care in a tech lead role. I’m 100% sure that any of those companies would I have immediately questioned my judgment for suggesting DeepSeek if had been a thing.

Absolutely none of them would ever have touched DeepSeek.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#229

Earlier quoted context omitted.

So much worse for American companies. This only means that they will be uncompetitive with similar companies that use models with realistic costs.

I can’t think of a single major US company that is big internationally that is competing on price.

Any car company. Uber.

All tech companies offering free services.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#230

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> Infrastructure owners with access to the cheapest energy will be the long run winners in AI. For a sufficiently low cost to orbit that may well be found in space, giving Musk a rather large lead. By his posts he's currently obsessed with building AI satellite factories on the moon, the better to climb the Kardashev scale.

The performance bottleneck for space based computers is heat dissipation.

Earth based computers benefit from the existence of an atmosphere to pull cold air in from and send hot air out to.

A space data center would need to entirely rely on city sized heat sink fins.

Post reply on HN