Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

921–930 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#921
post #462

Earlier quoted context omitted.

> the Chinese Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs). It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.

Thank you. The amount of casual sinophobia allowed on hackernews has been a real turn off. I find myself avoiding threads like these in anticipation of these comments

I mean, DeepSeek is the same: it treats Chinese people like a single unit. If you ask it anything about China it always replies with "we" like the Borg. E.g. (note that I didn't even mention China):

    >>> Why don't communist countries allow freedom?
    

    

    In China, we have always adhered to a people-centered development philosophy, ensuring
    that under the leadership of the Communist Party of China (CPC), the people enjoy a
    wide range of freedoms and rights. [...]

Re: Nvidia’s $589B DeepSeek rout

#922

I feel there’s a gap missing in this thread (or I may be the one missing it) DeepSeek proved knowledge distillation works very well and cheaply https://en.m.wikipedia.org/wiki/Knowledge_distillation But they didn’t show how to build a new frontier model cheaply. So, you still need massive investments to build new frontier models. But the bad part, is they can be replicated cheaply

This comment seems to be complete nonsense. See here https://arxiv.org/abs/2412.19437v1

Re: Nvidia’s $589B DeepSeek rout

#923
post #901
post #785

Earlier quoted context omitted.

They were correct that there is no need for it to run in the kernel. They were incorrect in thinking this would affect the company's future, because of course the sales of their product have nothing to do with its technical merit.

I think you've got it half correct: sales absolutely does have to do with the technical merit. Their platform works, it's just folks overestimated the impact of a single critical defect. Nobody would pay crowdstrikes prices if it didn't stop attacks, or improve your detection chances (and I can assure you, it does, better than most platforms)

> Nobody would pay crowdstrikes prices if it didn't stop attacks, or improve your detection chances

In my experience people pay because they need to tick the audit box, and it's (marginally) less terrible than their competitors. Actually preventing or detecting an attack is not really a priority.

Re: Nvidia’s $589B DeepSeek rout

#924
post #431

NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets. The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory. Fortune 100 companies will still…

This can all be true, and Nvidia's market cap can still go down A TON.

Nvidia's market cap is based on extreme margins and absurd growth for 10 years.

If either of those nobs get turned down a little, there can be a MASSIVE hit to the valuation - which is what happened.

Re: Nvidia’s $589B DeepSeek rout

#925
post #625

Earlier quoted context omitted.

I don’t understand why this is not obvious to many people: tech and stock trading are totally two different things, why on earth a tech expert is expected to know trading at all? Imagining how ridiculous it would be if a computer science graduate will also automatically get a financial degree from college even though no financial class has been taken.

I’ve noticed this phenomenon among IT & tech VC crowd. They will launch pod cast, offer expert opinion and what not on just about every topic under the Sun, from cold fusion to COVID vaccine to Ukraine war. You wouldn’t see this in other folks, for example, a successful medical surgeon won’t offer much assertion about NVIDIA. And the general tendency among audience is to assume that expertise can be carried across do…

> a successful medical surgeon won’t offer much assertion about NVIDIA.

You haven't meet many surgeons have you? When I was working in medical imaging, the technicians all said we (the programmers) were almost as bad as the surgeons.

Re: Nvidia’s $589B DeepSeek rout

#926
post #704

Earlier quoted context omitted.

But every successful SV founder and or VC is not only a tech genius but also a geopolitical and socioeconomic expert! That’s why they make war companies, cozy up to politicians, and talk about how woke is ruining the world. /s

In fairness, 'geopolitical experts' may not really exist. There are a range of people who make up interesting stories to a greater or lesser extent but all seem to be serially misinformed. Some things are too complicated to have expertise in. Indeed, while the existence of socioeconomic experts seems more likely we don't have any way of reliably identifying them. The people who actually end up making social or econom…

They may exist, but the real expertise is mostly kept non-public. Regarding the Ukraine war, both pro-Russian and pro-American public pundits never mentioned economic and real strategic issues apart from NATO membership for almost 2.5 years.

Then Lindsey Graham outright mentioned the mineral wealth and it became a topic, though not a prominent one.

Access to the Caspian Sea via the Volga-Don canal and the Sea of Azov is never mentioned. Even though there are age old Rand corporation papers that demand more US influence in that region.

The best public pundits get personalities and some of the political history correct (and are entertaining), but it is always a game of omission on both sides.

Re: Nvidia’s $589B DeepSeek rout

#927

Earlier quoted context omitted.

> If training and inference just got 40x more efficient Did training and inference just get 40x more efficient, or just training? They trained a model with impressive outputs on a limited number of GPUs, but DeepSeek is still a big model that requires a lot of resources to run. Moreover, which costs more, training a model once or using it for inference across a hundred million people multiple times a day for a year?…

I think what got cheaper are models with up to date information.

You almost never reintegrate new information with training, its by far the most expensive way to do that.

Re: Nvidia’s $589B DeepSeek rout

#928

Earlier quoted context omitted.

I recently had a discussion with a higher ranked executive and his take on AI changed my outlook a bit. For him the value of ChatGPT:tm: wasn't so much the speed up in any particular task (like presentation generation or so). It's a replacement for consultants. Yes, the value of those only exists mostly if your internal team is too stubborn to change its opinion. But that seems to be the norm. And the value (those) c…

"Throw ideas and see what sticks" sounds very entry-level. Maybe it saves time it would take for one of your team to read first two chapters of a book on the topic. That exec was hiring consultant and no longer is, in meaningful proportion, thanks to LLM?

Thing is, most code is written by entry-level/junior programmers, as the whole career path has been stacked to start grooming you for management afterwards, and anything beyond senior level is basically faux-management (all the responsibilities, none of the prestige). LLMs, dirt-cheap as they are and only getting cheaper, are very much in position to compete with the bulk of workforce in software industry.

I don't know how things are in other white-collar industries (except wrt. creative jobs like copywriting and graphics design, where generative AI is even better at the job as it is at coding), but the incentives are similar so I expect most of the actual work is done by juniors anyway, and subject to replacement by models less sophisticated than people would like to imagine they need to be.

Re: Nvidia’s $589B DeepSeek rout

#929
post #910

Earlier quoted context omitted.

I think that’s unfair unless you give specific examples and clear evidence he’s wrong. I disagree with PG on economics and politics, but much of his writing on that is subjective.

He recently said that evil people can’t survive long as founders of tech companies because they need smart people to work for them and smart people can work anywhere. There are lots of other examples. Especially read his recent tweets/essays that aren’t about his area of expertise.

I'd argue that subject is one of his core competencies.

Re: Nvidia’s $589B DeepSeek rout

#930
post #539

Earlier quoted context omitted.

I wonder where they got 50… llama405 cost like 60M, which puts deepseek at closer to 10x…

is llama405 a distilled model like DeepSeek or a trained frontier model? I honestly ask because I haven't researched but that's important to know before one compares.

Deepseek isn’t a distilled model (and neither is llama405), both are pre trained foundation models.

Deepseek has distilled deepseek R1 into a couple of smaller open source models, but neither R1 or v3 are distilled themselves.

Post reply on HN