Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

361–370 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#361

Earlier quoted context omitted.

The success of Facebook basically proves that public brand perception does not matter at all

Facebook itself still has a big problem with it's lack of youth audience though. Zuck captured the boomers and older Gen X, which are the biggest demos of living people however.

> Zuck captured the boomers and older Gen X, which are the biggest demos of living people however.

In the developed world. I'm not sure about globally.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#362
post #312

Earlier quoted context omitted.

The performance bottleneck for space based computers is heat dissipation. Earth based computers benefit from the existence of an atmosphere to pull cold air in from and send hot air out to. A space data center would need to entirely rely on city sized heat sink fins.

And the presence of humans. Like with a lot of robotics, the devil is probably in the details. Very difficult to debug your robot factory while it's in orbit. That was fun to write but also I am generally on board with humanity pushing robotics further into space. I don't think an orbital AI datacentre makes much sense as your chips will be obsolete so quickly that the capex getting it all up there will be better spe…

Well, _if_ they can get launch costs down to 100 dollar / kg or so, the economics might make sense.

Radiative cooling is really annoying, but it's also an engineering problem with a straightforward solution, if mass-in-orbit becomes cheap enough.

The main reason I see for having datacentres in orbit would be if power in orbit becomes a lot cheaper than power on earth. Cheap enough to make up for the more expensive cooling and cheap enough to make up for the launch costs.

Otherwise, manufacturing in orbit might make sense for certain products. I heard there's some optical fibres with superior properties that you can only make in near zero g.

I don't see a sane way to beam power from space to earth directly.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#363
post #309

Earlier quoted context omitted.

For radiative cooling using aluminum, per 1000 watts at 300 kelvin: ~2.4m^2 area, ~4.8 liters volume, ~13kg weight. So a Starship (150k kg, re-usable) could carry about a megawatt of radiators per launch to LEO. And aluminum is abundant in the lunar crust.

We are jumping pretty far ahead for a planet that can barely put two humans up there, but it is a great deal of my scifi dreams in one technology tree so I'll happily watch them try.

The grandfather comment is perhaps mixing up two things:

If launch costs are cheap enough, you can bring aluminum up from earth.

But once your in-space economy is developed enough, you might want to tap the moon or asteroids for resources.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#364

Earlier quoted context omitted.

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

AirBnB is all in on DeepSeek and Qwen. https://sg.finance.yahoo.com/news/airbnb-picks-alibabas-qwen...

It's a customer service bot? And Airbnb is a vacation home booking site. It's pretty inconsequential

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#365

I am waiting for the first truly open model without any of the censorship built in. I wonder how long it will take and how quickly it will try to get shut down.

Most open models have been converted to uncensored versions. Search for the model name with the suffix "abliterated".

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#366
post #55

Earlier quoted context omitted.

You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k

and can be faster if you can get an MOE model of that

Deepseek is already a MoE

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#367

Earlier quoted context omitted.

>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…

That quote from Google is 2.5 years old.

I also cringed a bit about seeing a statement that old being cited, but all the events since then only proved google right, I'd say.

Improvements seem incremental and smaller. For all I care, I could still happily use sonnet 3.5.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#368
post #262

Earlier quoted context omitted.

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

The US companies are all basically GPUaas. I’m not sure what the financial model is here, but I like it.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#369
post #262

Earlier quoted context omitted.

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Short cheeky answer is that capitalists need to capture value and communists don’t. Less cheeky answer is that this is a good opportunity for China to make sure the world isn’t dominated by US-sourced AI models.

However in another way the US probably offers more free inference than China. What good is an open 600 billion parameter model to a poor person? A free account with ChatGPT might be more useful to them, though also more exploitative.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#370

Earlier quoted context omitted.

I wonder how well the opthalmologist is doing. These guys are going to be paying him a visit playing around with those lasers and no PPE.

Eh, I don't see the risk, no pun intended. It's not collimated, and it's not going to be in focus anywhere but on-target. It's also probably in the long-wave range >>1000 nm that's not focused by the eye. At the end of the day it's no different from any other source of spot heating. I get more nervous around some of the LED flashlights you can buy these days. I want one. Hot air blows.

It's 45w of lasing power. I have a scar on my hand that's 15 years old from running one of those at 10% power and getting a reflection from a bare metal sheet.

This will absolutely scar, if not char, your cornea faster than you can blink.

Post reply on HN