Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

261–270 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#261

Earlier quoted context omitted.

Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.

This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…

Ah, so exactly like Uber, Netflix, Microsoft, Amazon, Facebook and so on have done to the rest of the world over the last few decades then?

Where do you think they learnt this trick? Years lurking on HN and this post's comment section wins #1 on the American Hypocrisy chart. Unbelievable that even in the current US people can't recognize when they're looking in the mirror. But I guess you're disincentivized to do so when most of your net worth stems from exactly those companies and those practices.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#262
post #211

Earlier quoted context omitted.

Do we know why?

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#263

Earlier quoted context omitted.

Pure models clearly aren’t the monetizing strategy, use of them on existing monetized surfaces are the core value. Google would love a cheap hq model on its surfaces. That just helps Google.

Hmmm but external models can easily operate on any "surface". For instance Claude Code simply reads and edits files and runs in a terminal. Photo editing apps just need a photo supplied to them. I don't think there's much juice to squeeze out of deeply integrated AI as AI by its nature exists above the application layer, in the same way that we exist above the application layer as users.

Gemini is the most used model on the planet per request.

All the facts say otherwise to your thoughts here.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#264

Earlier quoted context omitted.

[flagged]

This is how crazy and nationalistic people are getting. I'm an American citizen, though I am critical of the US government, and have no allegiances to China. What do you think America is doing to every country, even allies (which has been highly publicized)? Why would a country being constantly attacked by American intelligence and propaganda not want to counter that? https://www.reuters.com/world/europe/us-security-…

[dead]

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#265
post #5

I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper

Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…

Qwen, Hunyuan, and WAN are three of the major competitors in the vision, text-to-image, and image-to-video spaces. They are quite competitive. Right now WAN is only behind Google's Veo in image-to-video rankings on llmarena for example

https://lmarena.ai/leaderboard/image-to-video

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#266
post #257
post #180

Earlier quoted context omitted.

No… Nobody I work for will touch these models. The fear is real that they have been poisoned or have some underlying bomb. Plus y’know, they’re produced by China, so they would never make it past a review board in most mega enterprises IME.

For what it's worth, this is complete insanity when practically every mega enterprises' hardware is largely Made in China.

Enterprise hardware isn’t the issue. It’s the software. How much enterprise hardware is running with Chinese software? The US basically bans any hardware with Chinese software that can disrupt infrastructure.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#267

Earlier quoted context omitted.

What Chinese built infrastructure tech where information can be exfiltrated or cause any real damage are American companies buying? Chinese communication tech is for the most part not allowed in any American technology.

80% of the parts in iPhones are manufactured in China, and they have completely and utterly dominated in Enterprise (Ever heard of someone using a Blackberry in 2025? Me neither.) so there’s one example.

The software is made by Apple. Hardware can’t magically intercept communications and the manufacturing is done mostly in Taiwan. If Apple doesn’t have a process to protect its operating system from supply chain attacks, it would be derelict

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#268
post #123
post #18

Earlier quoted context omitted.

There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…

These Americans have no comprehension of intelligence being used to benefit humanity instead of being used to fund a CEO's new yacht. I encourage them to visit China to see how far the USA lags behind.

Lags behind meaning we haven't covered our buildings in LEDs?

America is mostly suburbs and car sewers but that's because the voters like it that way.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#269
post #262

Earlier quoted context omitted.

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Over reliance on investors who demand profits more than engineering.

The best innovation always happens before being tainted by investment.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#270
post #208

Pretty amazing that a relatively small Chinese hedge fund can build AI better than almost anyone.

Yeah they've consistently delivered. At the same time there are persistent whispers that they're not all that small and scruffy as portrayed either.

Anthropic also said their development costs aren't very different.
Post reply on HN