Live data from Hacker News

Nvidia's Risky Business

stratechery.com

151–160 of 185 posts

Re: Nvidia's Risky Business

#151

Nvidia's biggest advantage in AI has never been only their hardware performance but how entrenched their software is in ML research that flowed down stream. However, if you've actually used CUDA C/C++, it's pretty one of the worst software development ecosystem imaginable: you get all the footgun of regular C++, plus GPU compute pretending to be C++ and but doesn't actually behave like C++ because CPU and GPU compute…

That's really interesting. I have no experience writing anything that involves GPUs/TPUs, but over the years I've consistently read that CUDA is the "real moat" of Nvidia, which I never totally believed, but the way you describe makes it seem like it's not actually a moat in the slightest. It just happens to be an ecosystem associated with hardware that is not only considered the gold standard but happens to be more…

CUDA has many problems, but I would say less problems than ROCm, etc; even before considering the ecosystem and that more people (or open source projects) have already solved CUDA's problems for you.

Ironically, there was an open source project that was making great progress on CUDA compatibility on AMD hardware. AMD hired the lead developer, and then he shut down the project.

Re: Nvidia's Risky Business

#152

Nvidia's biggest advantage in AI has never been only their hardware performance but how entrenched their software is in ML research that flowed down stream. However, if you've actually used CUDA C/C++, it's pretty one of the worst software development ecosystem imaginable: you get all the footgun of regular C++, plus GPU compute pretending to be C++ and but doesn't actually behave like C++ because CPU and GPU compute…

The CUDA runtime coming with a gazillion reasonably decent kernels (DNN, BLAS, CUTLASS) and a concurrency system (NCCL) is a big deal; especially in the “early days” very few researchers or development runtimes were even writing their own kernels or dealing with CUDA C++ extensively, they were wrapping the ones NVidia gave them.

I do agree that it’s really not great, and I also have never been a strong believer in the CUDA moat overall; as the need for GPUs moves from research to production (inference), companies are plenty willing to build software from scratch anyway (and we see this with AMD GPUs being in plenty high demand in the datacenter and enthusiast market now).

Re: Nvidia's Risky Business

#153

For awhile I've found two things hard to square, that the hardware and software making up current gen AI will bring us to a socioeconomic singularity, and the reality the thing they're mostly trying to emulate is a few pounds of meat and fat running on tens of watts equivalent. On one hand the current AIs are obviously super human in some tasks, get completely dunked on in others by far simpler organisms. My cat can…

> a few pounds of … running on tens of watts equivalent.

Aren’t you already describing a laptop computer? Add a local LLM and you’re fully there

Re: Nvidia's Risky Business

#154

Earlier quoted context omitted.

I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side. The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have…

> efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute I buy that. Jevon's Paradox, sure. > and we are not going to run out of economically useful things to do with it anytime soon on the demand side This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there…

Replace "chat-bot" with "Human Being" because the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example.

Now for every human replacement, that is 1 unit less of communication and bureaucratic burden (HR, middle management etc) that the org requires.

Re: Nvidia's Risky Business

#155
post #151

Earlier quoted context omitted.

That's really interesting. I have no experience writing anything that involves GPUs/TPUs, but over the years I've consistently read that CUDA is the "real moat" of Nvidia, which I never totally believed, but the way you describe makes it seem like it's not actually a moat in the slightest. It just happens to be an ecosystem associated with hardware that is not only considered the gold standard but happens to be more…

CUDA has many problems, but I would say less problems than ROCm, etc; even before considering the ecosystem and that more people (or open source projects) have already solved CUDA's problems for you. Ironically, there was an open source project that was making great progress on CUDA compatibility on AMD hardware. AMD hired the lead developer, and then he shut down the project.

ZLUDA is still alive. AMD sponsored it and the project was briefly halted during a dispute with them, but it’s been making steady progress.

It doesn’t really make sense for AMD themselves or most use cases, though; any compatibility shim just adds problems on top of problems, and for AMD, entrenching a competitors technology even more never really seemed like a great idea.

Re: Nvidia's Risky Business

#156

Earlier quoted context omitted.

That's really interesting. I have no experience writing anything that involves GPUs/TPUs, but over the years I've consistently read that CUDA is the "real moat" of Nvidia, which I never totally believed, but the way you describe makes it seem like it's not actually a moat in the slightest. It just happens to be an ecosystem associated with hardware that is not only considered the gold standard but happens to be more…

Retraining a large high paid user base is often a non-starter. To put this in perspective, Boeing’s eventual retraining costs for all the pilots for the 737Max was around 5 billion dollars. Looking at software more specifically the Linux foundation reported based on software dev salaries in 2008 it would be 1.4 billion to only write the Linux kernel. Up until about 2023 there wasn’t enough money involved to have any…

As agentic coding continues to improve, won't it get easier for devs to retrain for new languages/frameworks/etc.?

Re: Nvidia's Risky Business

#157

Earlier quoted context omitted.

> efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute I buy that. Jevon's Paradox, sure. > and we are not going to run out of economically useful things to do with it anytime soon on the demand side This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there…

Replace "chat-bot" with "Human Being" because the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example. Now for every human replacement, that is 1 unit less of communication and bureaucratic burden (HR, middle management etc) that the org requires.

> the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example.

I guess I don't know what to say except that my experience is the polar opposite of yours.

I moved to a new state at the beginning of the year. Needed a new doctor, needed to schedule apartment tours, needed to talk to my employer about insurance and relocation stuff, etc etc. Lots of chatbots, a handful of humans. Humans consistently did what I needed them to do, the chatbots just didn't. I could list examples but I'd be typing all night.

And yknow what, my one call with Comcast to get my internet set up was downright pleasant. The rep was knowledgeable and a good conversationalist.

Re: Nvidia's Risky Business

#158
post #17

In many investment theses - like Nvidia's bet that demand for compute will keep growing - the first order assumption is usually correct. Yes, demand for more compute, chips, infrastructure is huge and each year some additional data centers will be built. Where such investment bets usually fail is in the second-order assumptions: Ie. the expectation of the growth of demand. This is where there's a high chance that the…

What makes this insanely hard to predict is that the compute needed for the same quality output has roughly gone down 90% every 18 months for ~5 years. 1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models). 2) We don't know when the appetite for higher cost models might go down and…

  It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
I doubt it. If LLM inference efficiency is 50x better than today, then there could be 1000x increase in inference volume due and we'll end up needing even more chips.

Jevons paradox should win out for a long time for AI.

When internet connections got faster than 56k modems, we didn't use the same amount of bandwidth but faster. We used more bandwidth doing things like 4k streaming. I see the same in AI inference. If AI inference is that much more efficient, it will just enable more use cases for AI.

See for example, internet traffic over time: https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS779VS...

Even after so many years, internet traffic continues to grow at an increasing rate.

Re: Nvidia's Risky Business

#159
post #145
post #37

[flagged]

Please don't post grouchy comments like this on HN. The guidelines make it clear we're trying for something better here. https://news.ycombinator.com/newsguidelines.html

I did not mean to be snarky. But still, the proposition that a railroad bankruptcy “led to world war” (direct quote from the article) 40 years later is ridiculous.

Re: Nvidia's Risky Business

#160
post #69

Earlier quoted context omitted.

When efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.

Very few consumers are going to spend multiple thousands of dollars to save $10 per month. Companies absolutely will to save hundreds per month per employee, but that's not consumer hardware.

Well if the trend that the comment further up in this thread claimed continues and compute requirements keep dropping exponentially then perhaps in a few years you can have today’s frontier performance on the normal laptop you already have on your desk anyway.
Post reply on HN