Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

191–200 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#191
post #85

Earlier quoted context omitted.

I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

AirBnB is all in on DeepSeek and Qwen.

https://sg.finance.yahoo.com/news/airbnb-picks-alibabas-qwen...

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#192
post #189

Earlier quoted context omitted.

>winning on cost-effectiveness Nobody is winning in this area until these things run in full on single graphics cards. Which is sufficient compute to run even most of the complex tasks.

I mean, there are lots of models that run on home graphics cards. I'm having trouble finding reliable requirements for this new version, but V3 (from February) has a 32B parameter model that runs on "16GB or more" of VRAM[1], which is very doable for professionals in the first world. Quantization can also help immensely. Of course, the smaller models aren't as good at complex reasoning as the bigger ones, but that se…

> but V3 (from February) has a 32B parameter model that runs on "16GB or more" of VRAM[1]

No. They released a distilled version of R1 based on a Qwen 32b model. This is not V3, and it's not remotely close to R1 or V3.2.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#193
post #85

Earlier quoted context omitted.

As much I agree with your sentiment, but I doubt the intention is singular.

I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…

[dead]

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#194

Earlier quoted context omitted.

[flagged]

Should I root for the democratic OpenAI, Google or Microsoft instead?

Further more, who thinks our little voices matter anymore in the US when it comes to the investor classes?

And if they did, having a counterweight against corrupt self-centered US oligarchs/CEOs is actually one of the biggest proponents for an actual powerful communist or other model world power. The US had some of the most progressive tax policies in its existence when it was under existential threat during the height of the USSR, and when their powered started to diminish, so too did those tax policies.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#195
post #14

Earlier quoted context omitted.

People with basement rigs generally aren't the target audience for these gigantic models. You'd get much better results out of an MoE model like Qwen3's A3B/A22B weights, if you're running a homelab setup.

Yeah I think the advantage of OSS models is that you can get your pick of providers and aren't locked into just Anthropic or just OpenAI.

Reproducibility of results are also important in some cases.

There are consumer-ish hardware that can run large models like DeepSeek 3.x slowly. If you're using LLMs for a specific purpose that is well-served by a particular model, you don't want to risk AI companies deprecating it in a couple months and push you to a newer model (that may or may not work better in your situation).

And even if the AI service providers nominally use the same model, you might have cases where reproducibility requires you use the same inference software or even hardware to maintain high reproducibility of the results.

If you're just using OpenAI or Anthropic you just don't get that level of control.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#196
post #85

Earlier quoted context omitted.

I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

really a testament to how easily the us govt has spun a china bad narrative even though it is mostly fiction and american exceptionalism

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#197

To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.

the motive is to prevent us dominance of this space, which is a good thing

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#198

Earlier quoted context omitted.

I was just watching this video about a Chinese piece of industrial equipment, designed for replacing BGA chips such as flash or RAM with a good deal of precision: https://www.youtube.com/watch?v=zwHqO1mnMsA I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.

I wonder how well the opthalmologist is doing. These guys are going to be paying him a visit playing around with those lasers and no PPE.

Eh, I don't see the risk, no pun intended. It's not collimated, and it's not going to be in focus anywhere but on-target. It's also probably in the long-wave range >>1000 nm that's not focused by the eye. At the end of the day it's no different from any other source of spot heating. I get more nervous around some of the LED flashlights you can buy these days.

I want one. Hot air blows.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#199

Earlier quoted context omitted.

[flagged]

> CrowdStrike researchers next prompted DeepSeek-R1 to build a web application for a Uyghur community center. The result was a complete web application with password hashing and an admin panel, but with authentication completely omitted, leaving the entire system publicly accessible. > When the identical request was resubmitted for a neutral context and location, the security flaws disappeared. Authentication checks…

not convincing. have you tried saying "free palestine" on a college campus recently?
Post reply on HN