Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

961–970 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#961
post #666

Earlier quoted context omitted.

Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…

We don't know for example what a larger model can do with the new techniques DeepSeek is using for improving/refining it. It's possible the new models on their [own] failed to show progress but a combination of techniques will enable that barrier to be crossed. We also don't know what the next discovery/breakthrough will be like. The reward for getting smarter AI is still huge and so the investment will likely remain…

Pending me getting an understanding of what those advances were, maybe?

But making things smaller is different than making them more powerful, those are different categories of advancement.

If you've noticed, models of varying sizes seem to converge on a narrow window of capabilities even when separated by years of supposed advancement. This should probably raise red flags

Re: Nvidia’s $589B DeepSeek rout

#962

Earlier quoted context omitted.

Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…

> You can't do the distill/magnify cycle like you do with alphago are you sure? people are saying that there’s an analogous cycle where you use o1-style reasoning to produce better inputs to the next training round

KIND OF

if you've tried to get o1 to give you outputs in a specific format, it often just tells you to take a hike. It's a stubborn model, which implies a lot

This is speculation, but it seems that the main benefit of reasoning models is that they provide a dimension along which RL can be applied to make them better at math and maybe coding, things with verifiable outputs.

Reasoning models likely don't learn better reasoning from their hidden reasoning tokens, they're 1) trying to find a magic token which when raised to its attention make it more effective (basically give it room to say something that jogs its memory) or 2) it is trying to find a series of steps which do a better job of solving a specific class of problem than a single pass does, making it more flexible in some senses but more stubborn along others

Reasoning data as training data is a poison pill, in all likelihood, and just makes a small window of RL vulnerable problems easier to answer (when we have systems that don't better). It doesn't really plan well, doesn't truly learn reasoning, etc

Maybe seeing the actual output of o3 will change my mind but I'm horrifically bearish on reasoning models

Re: Nvidia’s $589B DeepSeek rout

#963

Earlier quoted context omitted.

I feel like this is a symptom of our broken economic system that has allowed too much cash to be trapped in the markets, forever making mostly imaginary numbers go up while the middle class gets squeezed and the poor continue to suffer. A fundamental feature of a capitalist system you can use money to make more money. That's great for growing wealth. But you have to be careful, it's like a sound system at a concert.…

Or you could, you know, like just stop printing money and the problem goes away.

How does not printing money make any difference here? It's not like people put trillions of cash into nVidia. The cap is just outstanding shares times what they were last sold for, if someone buys at a higher price suddenly lots of phantom money appears and everybody who owns shares gets richer on paper.

Re: Nvidia’s $589B DeepSeek rout

#964

Earlier quoted context omitted.

This is partially why Apple is the one that stands to gain more, and it showed. Their "small models, on device" approach can only be perfected with something like DeepSeek, and they're not exposed to NVIDIA pricing, nor have to prove investors that their approach is still valid.

Until AGI removes the need for iOS. Apple is not immune to AI disruption. The Rabbit R1 was a scam but the concept was the right approach. It was just 5 years too early. You don’t need an iPhone with AGI. You just need a 5G device with a screen, connected to an AGI.

Perennial reminder that we do not have any real evidence that we are anywhere close to AGI, or that "throwing more resources at LLMs" is even theoretically a possible way to get to an AGI.

"Lots of people with either a financial motivation to say so or a deep desire for AGI to be real Soon™ said they can do it" is not actual evidence.

We do not know how to make an AGI. We do not know how to define an AGI. It is hypothetically possible that we could accidentally stumble into one, but nothing has actually shown that, and counting on it is a fool's bet.

Re: Nvidia’s $589B DeepSeek rout

#965

Earlier quoted context omitted.

Probably not. If the price of Nvidia is dropping, it's because investors see a world where Nvidia hardware is less valuable, probably because it will be used less. You can't do the distill/magnify cycle like you do with alphago. LLM models have basically stalled in their base capabilities, pre training is basically over at this point, so the news arms race will be over marginal capability gains and (mostly) making th…

This argument ignores scaling laws

It really doesn't lol. Those laws are like Moore's law, an observation rather than something Fundamental like laws in physics

The scaling has been plateauing, and half that equation is quality training data which is totally out at this point.

Maybe reasoning models will help produce synthetic data but that's still to be seen. So far the only benefit reasoning seems to bring is fossilizing the models and improving outputs along a narrow band of verifiable answers that you can do RL on to get correct

Synthetic data maybe buys you time, but it's one turn of the crank and not much more

Re: Nvidia’s $589B DeepSeek rout

#966
post #625

Earlier quoted context omitted.

I don’t understand why this is not obvious to many people: tech and stock trading are totally two different things, why on earth a tech expert is expected to know trading at all? Imagining how ridiculous it would be if a computer science graduate will also automatically get a financial degree from college even though no financial class has been taken.

I’ve noticed this phenomenon among IT & tech VC crowd. They will launch pod cast, offer expert opinion and what not on just about every topic under the Sun, from cold fusion to COVID vaccine to Ukraine war. You wouldn’t see this in other folks, for example, a successful medical surgeon won’t offer much assertion about NVIDIA. And the general tendency among audience is to assume that expertise can be carried across do…

This is exacerbated by the tendency in popular media to depict a Scientist character, who can do all kinds of Science (which includes technology, all kinds of computing, and math).

Re: Nvidia’s $589B DeepSeek rout

#967

Earlier quoted context omitted.

With the crowdstrike outage earlier last year it was incredible how many hidden security and kernel "experts" came out crawling from the woodwork, questioning why anything needs to run in the kernel and predicting the company's demise.

the experts were correct. in 2024 there are now OS APIs that provide the same observability and control with much less risk involved.

Pretty sure every vendor uses the kernel. There must be a reason behind none of them exclusively using the APIs.

Re: Nvidia’s $589B DeepSeek rout

#968
post #927

Earlier quoted context omitted.

I think what got cheaper are models with up to date information.

You almost never reintegrate new information with training, its by far the most expensive way to do that.

...and that got cheaper? Not sure your point.

At some point, the models _have_ to do "continuous integration" to provide the "AGI" that's wanted out of this tech.

Re: Nvidia’s $589B DeepSeek rout

#969

Earlier quoted context omitted.

Jevon's paradox ony applies under certain conditions. It remains to be seen if it will hold in this case.

What do you understand those conditions to be?

Output quantity consumed (almost) always increases with falling inputs (ie, costs, whether in dollars or GPUs). But for Jevon's paradox to hold, the slope of quantity-consumption-increase-per-falling-costs must exceed a certain threshold. Otherwise, the result is just that quantity consumed increases while quantity of inputs consumed decreases.

Applied to AI and NVIDIA, the result of an increase in the AI-per-GPU on demand for GPUs depends on the demand curve for AI. If the quantity of AI consumed is completely independent of its price, then the result of better efficiency is cheaper AI, no change in AI quantity consumed, and a decrease in the number of GPUs needed. Of course, that's not a realistic scenario.

(I'm using "consumed" as shorthand; we both know that training AIs does not consume GPUs and AIs are also not consumed like apples. I'm using "consumed" rather than the term "demand" because demand has multiple meanings, referring both to a quantity demanded and a bid price, and this would confuse the conversation).

But a scenario that is potentially realistic is that as the efficiency of training/serving AI drops by 90%, the quantity of AI consumed increases by a factor of 5, and the end result is the economy still only needs half as many GPUs as it needed before.

For Jevons paradox to hold, if the efficiency of converting GPUs to AI increases by X, resulting in a decrease in price by 1/X, the quantity of AI consumed must increase by a factor of more than X as a result of that price decrease. That's certainly possible, but it's not guaranteed; we basically have to wait to observe it empirically.

There's also another complication: as the efficiency of producing AI improves, substitutes for datacenter GPUs may become viable. It may be that the total amount of compute hardware required to train and run all this new AI does increase, but big-iron datacenter investments could still be obsoleted by this change because demand shifts to alternative providers that weren't viable when efficiency was low. For example, training or running AIs on smaller clusters or even on mobile devices.

If tech CEOs really believe in Jevons Paradox, it means that last month when they decided to invest $500 billion in GPUs, then this month after learning of DeepSeek, they now realize $500 billion is not enough and they'll need to buy even more GPUs, and pay even more each one. And, well, maybe that's the case. There's no doubt that demand for AI is going to keep growing. But at some point, investment in more GPUs trades off against other investments that are also needed, and the thing the economy is most urgently lacking ceases to be AI.

Re: Nvidia’s $589B DeepSeek rout

#970

Earlier quoted context omitted.

ChatGPT has over $10 million paying subscriber. No I am not counting the people using the API programmatically

And they're still burning billions with no end in sight.

Are they losing billions on training or inference? If their current products - ChatGPT and the API - are profitable ie the inference cost is less than they charge, they have a long term sustainable business.
Post reply on HN