Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

461–470 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#461

Earlier quoted context omitted.

It's 45w of lasing power. I have a scar on my hand that's 15 years old from running one of those at 10% power and getting a reflection from a bare metal sheet. This will absolutely scar, if not char, your cornea faster than you can blink.

That's (again) less energy than a flashlight puts out these days, so the beam had to be tightly focused in your case. That isn't how these things work. There is nothing special about "lasing power." It amounts to a 45-watt light bulb, nothing more and nothing less.

A 45 watt light bulb spreads the energy in all directions - at 1 meter away that's about 3 watts in every square meter or roughly 0.000003 watts per square millimeter. The laser is putting 45 watts into that same square millimeter at the same distance.

Of course the laser is tightly focused. That's pretty much one of the defining properties of laser devices. How else do you think the laser is heating the microprocessors in the video?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#462
post #437

Earlier quoted context omitted.

I worked inside AWS consulting department for 3 years (AWS ProServe) and now I work as a staff consultant for a 3rd AWS partner. I have been on enough sales calls, seen enough go to market training materials and flown out to customers sites to know how these things work. AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it. Microsoft doesn’…

> AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it. Despite all that and whatever you say, the fact is you do compete. It doesn't have to be a race to the bottom. So Cloudfront free tier and the latest discount bundles etc aren't to compete? People have also negotiated private pricing way below list price and a lot cheaper than competit…

I am well aware that Netflix doesn’t pay the same price for AWS services that “Joe Bob’s Fish Tackle and WordPress shop”. All big companies give discounts to large companies as part of negotiations which is different from “we are the low cost leader”.

All technology gets cheaper over time. There is a difference between lowering price in response to competitors and finding the profit maximizing price based on supply and demand.

AWS was lowering prices to increase demand before GCP and Azure were a thing.

Jassy said right before he became CEO of Amazon and he was still over AWS that only 5% of IT spend was on any cloud provider. They are capturing non consumption and marketing value of AWS vs that.

While I don’t have any insider experience about Azure, looking on the outside, I would think that Azure’s go to market is also not competing against AWS on price, but trying to get on prem customers on Azure.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#463

Earlier quoted context omitted.

That's (again) less energy than a flashlight puts out these days, so the beam had to be tightly focused in your case. That isn't how these things work. There is nothing special about "lasing power." It amounts to a 45-watt light bulb, nothing more and nothing less.

A 45 watt light bulb spreads the energy in all directions - at 1 meter away that's about 3 watts in every square meter or roughly 0.000003 watts per square millimeter. The laser is putting 45 watts into that same square millimeter at the same distance. Of course the laser is tightly focused. That's pretty much one of the defining properties of laser devices. How else do you think the laser is heating the microprocess…

They will be using a beam spreader to conform to the size of the targeted IC, which is usually on the order of 5x5 mm and up. For smaller parts they will be reducing the power.

They shouldn't be focusing it to a point under any conditions. Whether it's as safe as it could be is a different question, of course. For instance, you'd like to think that the act of configuring it for a smaller beam footprint would reduce the power at the same time, as opposed to requiring a separate adjustment that might be overlooked by the operator. Would have been nice if the video had addressed that and other safety considerations, for sure.

A lot depends on the exact wavelength. 1400 nm and longer is much less worrisome than near-visible IR.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#464
post #178

Earlier quoted context omitted.

that's not really a good summary of what MoEs are. you can more consider it like sublayers that get routed through (like how the brain only lights up certain pathways) rather than actual separate models.

The gains from MoE is that you can have a large model that's efficient, it lets you decouple #params and computation cost. I don't see how anthropomorphizing MoE brain affords insight deeper than 'less activity means less energy used'. These are totally different systems, IMO this shallow comparison muddies the water and does a disservice to each field of study. There's been loads of research showing there's redundan…

> I don't see how anthropomorphizing MoE brain affords insight deeper than 'less activity means less energy used'.

I'm not saying it is a perfect analogy, but it is by far the most familiar one for people to describe what sparse activation means. I'm no big fan of over-reliance on biological metaphor in this field, but I think this is skewing a bit on the pedantic side.

re: your second comment about pruning, not to get in the weeds but I think there have been a few unique cases where people did lose some of their brain and the brain essentially routed around it.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#465

To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.

the motive is to prevent us dominance of this space, which is a good thing

And the next question is what have they some with power historically, and what are they liable to do in the future with said power. Limiting scope to AI is shortsighted and doesn't speak to the concerns people have beyond an Ai Race

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#466

Earlier quoted context omitted.

the motive is to prevent us dominance of this space, which is a good thing

And the next question is what have they some with power historically, and what are they liable to do in the future with said power. Limiting scope to AI is shortsighted and doesn't speak to the concerns people have beyond an Ai Race

It's a fair question, but my view of America's influence on world affairs has been dismal. China by contrast has not had a history of invading its neighbors, though I strongly criticize their involvement in the American attack on Cambodia and Vietnam (China supported the Khmer Rouge and briefly invaded Vietnam but was quickly pushed back, a reason Mao is sometimes criticized as having a good early period and a bad late period).

Meanwhile, America has been causing death and destruction around the world. It's easy to make lists: Vietnam, Iraq, Gaza, Cuba, South and Central America etc etc.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#467
post #412

Earlier quoted context omitted.

cerebras AI offers models at 50x the speed of sonnet?

if that's an honest question, the answer is pretty much yes, depending on model.

the question mark was expressing confusion.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#468

Earlier quoted context omitted.

Yes, he did, and it was fundamental to his entire economic philosophy: https://en.wikipedia.org/wiki/Tendency_of_the_rate_of_profit...

I'm not seeing anywhere in that page anything about an assumed static basket of human wants and needs. Maybe I missed it -- can you point out where you saw that? Interesting, though, that per the very same article someone like Adam Smith concurred empirically with Marx's observation on the titular tendency of rates of profit to fall. This suggests to me it likely had some meat to it.

Without going too deep on it (I used to be a fan in university as a silly youth), the tendency of the rate of profit to fall is the key aspect of Marx's crisis theory.

Basically dude thought the competition inherent in capitalism would cause all profit to be competed to zero leading to an eventual 'crisis' and collapse of the capitalist means of production.

Implicit in this assumption is the idea that the things humans need and want changes/evolves in a predictable way, and not in a chaotic/fractal/reflexive way (which is what actually happens).

An eventual static basket of desired goods would be the only mechanism by which competition ever could compete profits to zero. If the basket is dynamic/reflexive/evolving, there's constantly new gaps opening between human desires and market offerings to arbitrage for profit. You can just look at the average profit margins of S&P500 companies over time to see they are not falling.

The further we get from subsistence worries (Adam Smith's invisible hand has pulled virtually the entire globe out of living in the dirt), the more divergent and higher abstraction these wants and needs become, and hence the profit opportunities are only increasing -- which is how the economy grows (no, it's not a fixed pie, another Marxian fallacy).

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#469

Earlier quoted context omitted.

A 45 watt light bulb spreads the energy in all directions - at 1 meter away that's about 3 watts in every square meter or roughly 0.000003 watts per square millimeter. The laser is putting 45 watts into that same square millimeter at the same distance. Of course the laser is tightly focused. That's pretty much one of the defining properties of laser devices. How else do you think the laser is heating the microprocess…

They will be using a beam spreader to conform to the size of the targeted IC, which is usually on the order of 5x5 mm and up. For smaller parts they will be reducing the power. They shouldn't be focusing it to a point under any conditions. Whether it's as safe as it could be is a different question, of course. For instance, you'd like to think that the act of configuring it for a smaller beam footprint would reduce t…

OK, put you face in front of a 45w co2 laser tube and report your results.

The laser is collimated but not focused so by your logic it will be fine.

This is advice on par with eating tide pods.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#470
post #451
post #273

Earlier quoted context omitted.

Several of your comments in this subthread have broken the guidelines. The guidelines ask us not to use HN for political/ideological battle and to "assume good faith". They ask us to "be kind", "eschew flamebait", and ask that "comments should get more thoughtful and substantive, not less as a topic gets more divisive." The topic itself, like any topic, is fine to discuss here, but care must be taken to discuss it in…

Can you, by any chance, delete my account? I have tried to do so before but it is not possible through the GUI. And I see you are associated with HN. Other than that let's be very clear that there was no personal attack. You left out the part where I explain why I think the comment was made in bad faith. I.e. the part that makes it not a personal attack. And a part which I, upon request, elaborated on in the same com…

We can disable your account if you email hn@ycombinator.com. That's in the FAQ – https://news.ycombinator.com/newsfaq.html.

And yes I am a moderator and it's my role to prevent flamewars and to encourage everyone to raise the standard of discourse here. In my comment I was trying to convey that multiple comments of yours were crossing too far into political battle and personal attack, and here are the main instances:

> That is just objectively incorrect, and fundamentally misunderstanding the basics of statehood

This counts as a personal swipe, and as fulminating.

> It is however entirely reasonable to assume that the comment I replied to was made entirely in bad faith

People can be mistaken or wrong, or just of a different opinion/assessment, without acting “entirely in bad faith”.

> "Baselessly" - I'm sorry but realpolitik is plenty of basis. China is a geopolitical adversary of both the EU and the US. And China will be the first to admit this, btw.

This is phrased in a snarky way.

The points you've made are fine to make, but the way you make them matters. Snarkiness, swipes, put-downs, accusations of bad faith (giving your reason "why" you think it was in bad faith doesn't make it OK) are all clearly against the guidelines.

I can accept that you didn't mean to break the guidelines, which is why I've politely asked you to familiarise yourself with them and try harder to follow them in future. It's a request not a scolding. It's not necessary to announce you want to quit HN in protest. (Though of course, eventually we would rather people leave if they prefer not to follow the guidelines.) Just making an effort to respect the guidelines and the HN community would be great.

Post reply on HN