Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

231–240 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#231
post #229

Earlier quoted context omitted.

I can’t think of a single major US company that is big internationally that is competing on price.

Any car company. Uber. All tech companies offering free services.

Is a “cheaper” service going to come along and upend Google or Facebook?

I’m not saying this to insult the technical capabilities of Uber. But it doesn’t have the economics that most tech companies have - high fixed costs and very low marginal costs. Uber has high marginal costs saving a little on inference isn’t going to make a difference.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#232
post #229

Earlier quoted context omitted.

I can’t think of a single major US company that is big internationally that is competing on price.

Any car company. Uber. All tech companies offering free services.

What American car company competes overseas on price?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#233

Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??

128 GB should be enough for anybody (just kidding). I hope the M5 Max will have higher RAM limits

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#234

To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.

Do you think it is from goodness of their heart that corporates support open source? E.g. Microsoft - VSCode and Typescript, Meta - PyTorch and React, Google - Chromium and Go.

Yet, we (developers, users, human civilization), benefit from that.

So yes, I cherish when Chinese companies release open source LLMs. Be it as it fits their business model (the same way as US companies) or from grants (the same way as a lot of EU-backed projects, e.g. Python, DuckDB, scikit-learn).

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#235

Earlier quoted context omitted.

Big enterprise with mostly private companies as their clients? Lol, yeah, that’s how they work from my personal experience. The reality is, if it’s not a tech-first enterprise and already outsource part of tech to a shop outside of NA (which is almost majority at this point), they will do absolutely everything to cut the costs.

I spent three years working in consulting mostly in public sector and education and the last two working with startups to mid size commercial interest and a couple of financial institutions. Before that I spent 6 years working between 3 companies in health care in a tech lead role. I’m 100% sure that any of those companies would I have immediately questioned my judgment for suggesting DeepSeek if had been a thing. Ab…

Why would you be presenting what AI tech you are using? You would tell them AI will come from Amazon using a variety of models.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#236

Earlier quoted context omitted.

> Infrastructure owners with access to the cheapest energy will be the long run winners in AI. For a sufficiently low cost to orbit that may well be found in space, giving Musk a rather large lead. By his posts he's currently obsessed with building AI satellite factories on the moon, the better to climb the Kardashev scale.

The performance bottleneck for space based computers is heat dissipation. Earth based computers benefit from the existence of an atmosphere to pull cold air in from and send hot air out to. A space data center would need to entirely rely on city sized heat sink fins.

For radiative cooling using aluminum, per 1000 watts at 300 kelvin: ~2.4m^2 area, ~4.8 liters volume, ~13kg weight. So a Starship (150k kg, re-usable) could carry about a megawatt of radiators per launch to LEO.

And aluminum is abundant in the lunar crust.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#237

Earlier quoted context omitted.

Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…

forgive me for bringing politics into it, are chinese LLM more prone to censorship bias than US ones ?

Yes extremely likely they are prone to censorship based on the training. Try running them with something like LM Studio locally and ask it questions the government is uncomfortable about. I originally thought the bias was in the GUI, but it's baked into the model itself.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#238

Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??

128 GB should be enough for anybody (just kidding). I hope the M5 Max will have higher RAM limits

M5 Max probably won’t, but M5 Ultra probably will

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#239
post #43

Earlier quoted context omitted.

Qwen Image and Image Edit were among the best image models until Nano Banana Pro came along. I have tried some open image models and can confirm , the Chinese models are easily the best or very close to the best, but right now the Google model is even better... we'll see if the Chinese catch up again.

I'd say Google still hasn't caught up on the smaller model side at all, but we've all been (rightfully) wowed enough by Pro to ignore that for now. Nano Banano Pro starts at 15 cents per image at Add in the power of fine-tuning on their open weight models and I don't know if China actually needs to catch up. I finetuned Qwen Image on 200 generations from Seedream 4.0 that were cleaned up with Nano Banana Pro, and got…

FWIW, Qwen Z-Image is much better than Seedream and people (redditors) are saying its better than Nano Banana in their first trials. Its also 7B I think, and open.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#240

Earlier quoted context omitted.

There are plenty of 3rd party and big cloud options to run these models by the hour or token. Big models really only work in that context, and that’s ok. Or you can get yourself an H100 rack and go nuts, but there is little downside to using a cloud provider on a per-token basis.

> There are plenty of 3rd party and big cloud options to run these models by the hour or token. Which ones? I wanted to try a large base model for automated literature (fine-tuned models are a lot worse at it) but I couldn't find a provider which makes this easy.

have you checked OpenRouter if they offer any providers who serve the model you need?
Post reply on HN