Live data from Hacker News

Nvidia’s $589B DeepSeek rout

finance.yahoo.com

991–1000 of 1001 posts

Re: Nvidia’s $589B DeepSeek rout

#991

Earlier quoted context omitted.

What do you understand those conditions to be?

Output quantity consumed (almost) always increases with falling inputs (ie, costs, whether in dollars or GPUs). But for Jevon's paradox to hold, the slope of quantity-consumption-increase-per-falling-costs must exceed a certain threshold. Otherwise, the result is just that quantity consumed increases while quantity of inputs consumed decreases. Applied to AI and NVIDIA, the result of an increase in the AI-per-GPU on…

Thanks for that.

There's a Jevons Paradox article up now and I'll put most of my thoughts there: https://news.ycombinator.com/item?id=42863808>

If you care to respond though, my first question would be what examples of falling input prices not subject to the Jevons Paradox are. Several of the more notorious ones involve energy, and that was Jevons's principle topic of study (The Coal Question most notably).

I've got my own theory of how technological mechanisms function, with an ontology of nine elements. Fuels are one of those, information is another. See prior comments: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...>

As might be pertinent to AI and LLM, whilst fuels and power applications seem to scale linearly against input (constant slope, if not 1:1 relation), information processing delivers far more variable returns, often with critical thresholds. Network effects and Metcalfe's Law are the best known of these (if highly inaccurate themselves, see Tilly-Odlyzko's refutation), but another is the limited returns of predictive and targeting applications.

For the latter, the 18 order of magnitude increase in computing power from 1965--2025 (60 years, about 20--30 Moore's Law cycles) has roughly doubled the length of accurate long-term weather forecasting from roughly 5 days to 10. It's made possible fully-resuable first-stage boosters for orbital spaceflight, which is visually impressive, but has only resulted in a five-fold reduction ($1,400/kg vs. $5,400/kg) in low-Earth orbit (LEO) launch costs (Falcon Heavy vs. Saturn V). SpaceX are looking for another factor of 2--4 reduction (to $250--600/kg), but that's still far less improvement than we've seen in raw compute. At some point orbital physics, the rocket equation, and fuel chemistry simply dominate other considerations.

Similarly, AdTech makes possible far more targeted advertising, but to heavily diminishing returns, the core result has been an abandonment of non-targetable media by advertisers, notably print and broadcast, as well as an arms-race between the browser (for a very small fraction of the market) and advertisers (the largest of which also has the largest browser marketshare), and a concentration of advertising revenue amongst two online entities, Google (a/k/a Alphabet) and Facebook (a/k/a Meta).

Which makes me wonder what applications AI LLMs might practically be put to. Advertising, manipulation, fraud, and propaganda certainly seem to be benefiting.

Re: Nvidia’s $589B DeepSeek rout

#992

Earlier quoted context omitted.

I agree hype is a big portion of it, but if DeepSeek really has found a way to train models just as good as frontier ones for a hundredth of the hardware investment, that is a substantial material difference for Nvidia's future earnings.

Is it? Training is only done once, inference requires GPUs to scale, especially for a 685B model. And now, there’s an open source o1 equivalent model that companies can run locally, which means that there’s a much bigger market for underutilized on-prem GPUs.

I'd be really curious about the split in hardware for training vs inference - I got the read that it was a very high ratio to the point the training is not a significant portion of the requisite hardware but instead the inference at scale sucks up most of the available datacenter gpu share.

Could be entirely wrong here - would love a fact-check by industry insider or journalist.

Re: Nvidia’s $589B DeepSeek rout

#993

Earlier quoted context omitted.

I agree hype is a big portion of it, but if DeepSeek really has found a way to train models just as good as frontier ones for a hundredth of the hardware investment, that is a substantial material difference for Nvidia's future earnings.

1. Nobody has replicated their DeepSeek's results on their reported budget yet. Scale.ai's Alexander Wang says they're lying and that they have a huge, clandestine H100 cluster. HuggingFace is assembling an effort to publicly duplicate the paper's claims. 2. Even if DeepSeek's budget claims are true, they trained their model on the outputs of an expensive foundation model built from a massive capital outlay. To truly…

https://xyzlabs.substack.com/p/berkeley-researchers-replicat...

Given they've reproduced earlier model's and vetted it - I think it's probably safe to assume that these new models are not out of thin air - but until somebody reproduces it, it's up in the air.

Re: Nvidia’s $589B DeepSeek rout

#994

Earlier quoted context omitted.

The other thing is that if this pushes the envelope further on what AI models can do given a certain hardware budget, this might actually change minds. The pushback against generative AI today is that much of it is deployed in ways that are ultimately useless and annoying at best, and that in turn is because the capabilities of those models are vastly oversold (including internally in companies that ship products wit…

An rag model can already sort your email. Its just that it costs too much to do that for the hoi polloi who think everything digital should be free forever.

It can do it if you're okay with hallucinations. Which is generally not the case.

Re: Nvidia’s $589B DeepSeek rout

#995

Earlier quoted context omitted.

I think that’s unfair unless you give specific examples and clear evidence he’s wrong. I disagree with PG on economics and politics, but much of his writing on that is subjective.

Happy to. Here's PG misrepresenting a wealth tax - https://nindalf.com/posts/wealth-tax/

Returning to this: even if he were somehow "misrepresenting" the idea of wealth tax - and there are many variants - that does not make him wrong.

My original point was that when PG is writing outside of his core competencies, it's usually about subjective opinions. Such as, is wealth tax a good or bad thing? That is a subject about which you can have an opinion, but there is no absolute.

My concern is that when people say "he doesn't know anything when he talks about, say, economics or politics" what they really mean is "I disagree with him and therefore he is wrong." (I disagree with PG on many such things, but that's like, just my opinion, man)

I guess I'm invested in the debate because I want people to be more open-minded and charitable and not less.

Re: Nvidia’s $589B DeepSeek rout

#996

Earlier quoted context omitted.

the experts were correct. in 2024 there are now OS APIs that provide the same observability and control with much less risk involved.

Pretty sure every vendor uses the kernel. There must be a reason behind none of them exclusively using the APIs.

yeah, it's money.

Re: Nvidia’s $589B DeepSeek rout

#997
post #711

Earlier quoted context omitted.

But did it not lead to net less electricity used for lighting?

With more efficiency offset by more light bulbs, electricity used for lighting has been roughly flat since 2010: https://www.iea.org/data-and-statistics/charts/global-electr... (For LLMs I wish that efficiency could lead to less electricity used for chips, but I think the best we can hope for is for electricity use to flatten out.)

From that graph it looks like residential went down quite a lot, which makes sense since that was the main usage of incandescents.

Re: Nvidia’s $589B DeepSeek rout

#998

90% of the comments in this thread make it clear that knowing about technology does not in any way qualify someone to think correctly about markets and equity valuations.

I think that’s correct, but it would be a more useful comment if you gave a few examples and explained the correct thinking.

lol I posted the 12,000 word article that caused the sell-off:

https://x.com/doodlestein/status/1884712920543621148?s=46

Re: Nvidia’s $589B DeepSeek rout

#999

Earlier quoted context omitted.

I recently had a discussion with a higher ranked executive and his take on AI changed my outlook a bit. For him the value of ChatGPT:tm: wasn't so much the speed up in any particular task (like presentation generation or so). It's a replacement for consultants. Yes, the value of those only exists mostly if your internal team is too stubborn to change its opinion. But that seems to be the norm. And the value (those) c…

"Throw ideas and see what sticks" sounds very entry-level. Maybe it saves time it would take for one of your team to read first two chapters of a book on the topic. That exec was hiring consultant and no longer is, in meaningful proportion, thanks to LLM?

The point isn't that the result from an LLM is particularly valuable. The point is that the advice that you get from your typical (management) consultant isn't particularly useful. And that's the only bar you need to clear.

There are basically two reasons for consultants:

Either you need some once-removed thing (be that you selling/buying some part to/from a competitor, accounting, lawyering). That part sensible people will not replace by LLMs. But here it isn't that you yourself lack the expertise at all. Here it is absolutely necessary, that someone else does the actual implementation.

Or you have some general "we need to do better" feeling. And here you have again two options: 1) you know _what_ your problem is and you just need the best solution there is. This is (obviously somewhat tongue-in-cheek) essentially corporate espionage. Again you cannot replace that with an LLM (or maybe you can I don't know), but you will pay a lot of money for it. Or 2) you don't know what the problem is. Now you are competing in finding an appropriate starting point with 20-somethings fresh from university that are not wanted in your organisation and therefore won't get access to the relevant information anyhow. So yeah, I'm willing to believe that the typical success rate of those consultancy projects is low to negative.

If you are only given a 3 week crash course in $BUSINESS, you won't be able to produce much more than a generic set of "have you thought about that?". And THAT is something that I believe LLMs to be reasonably good at. And they are dirt cheap and instantly available compared to any kind of human consultant.

Now I don't think that will necessarily a net-negative for consultants in general. I do think, similar to ATMs, that consultations are mostly becoming cheaper by that.

Re: Nvidia’s $589B DeepSeek rout

#1000

Earlier quoted context omitted.

This argument ignores scaling laws

It really doesn't lol. Those laws are like Moore's law, an observation rather than something Fundamental like laws in physics The scaling has been plateauing, and half that equation is quality training data which is totally out at this point. Maybe reasoning models will help produce synthetic data but that's still to be seen. So far the only benefit reasoning seems to bring is fossilizing the models and improving out…

They are derived from statistical laws unlike moores law

https://en.m.wikipedia.org/wiki/Chernoff_bound

I agree with you that they require data

Post reply on HN