Live data from Hacker News

GPU World

gpuworld.org

211–220 of 303 posts

Re: GPU World

#211
One of the amazing things is that when every has one GPU, they will actually have 1k-10k agents at their disposal.

LLMs and KV caches have amazing performance characteristics with concurrent throughput. It scales very non linearly. So the token throughput within a batch scales WAY faster than the tokens per second of each user.

This is the reason the LLM providers have such crazy margins on their costs.

Re: GPU World

#212
post #94

> Someday, such as in 2040, there may be available, for every human being, the performance equivalent of 'a B300 GPU for contemporary LLMs'. What would this world be like? If we talk about just LLMs, given how things have been going since ChatGPT, my bet it would not change that much. LLMs are not foundational technology such as Internet or Steam engine or Rail roads were. There are very few products that can build u…

I disagree with this take. While LLMs themselves are currently unreliable, the work done in the math community on hooking up creative LLMs to reliable verifiers like Lean show that it’s possible to construct systems where the unreliability is suppressed. For now, that still requires experts to set up and monitor, but I do believe that in a couple of decades we’ll make progress on how to do more mundane tasks in a rel…

I don't think you've said anything different than the person you're responding to. The difference primarily seems to be whether you think AI is transformative vs simply being another tool (which is really just a matter of perspective).

If you're not an expert software engineer, it's transformative. If you are, it's just another tool.

Re: GPU World

#213
post #133

Earlier quoted context omitted.

The TAM of electricity is 99.999% of humanity, yet there was 82 years between Alessandro Volta's Voltaic Pile experiments and the first electrical grid in the world, Pearl Street Station, New York City. A farmer may have seen an LLM answering their questions about how to make a less dry chicken sandwich, but he hasn't seen a specifically designed and tested harness that autonomously manages a crop harvester. There is…

Fair points (and I appreciate very much that they read like a human wrote them). As a counter point - and I know it's a weak one - a few years ago people were telling me that crypto was going to replace currency, and soon every person on the planet would be using it for every transaction. Many people believe that will still happen, but I don't think many serious people do. I want to believe there is big bucks in this…

> a few years ago people were telling me that crypto was going to replace currency, and soon every person on the planet would be using it for every transaction. Many people believe that will still happen, but I don't think many serious people do.

As you say, that is a weak counter point. I do get what you're saying, but crypto flaws were built in from the start:

- Deflationary currency isn't going to be spent, it is going to be hoarded, if it would have any value at all. The design was self defeating from the very start.

- Technical limits of PoW. Obvious right from the start, which led to the block wars. PoS "solves" that but with other costs. And it still doesn't solve the first point.

Despite these, crypto did get adopted, just not in any way that crypto maximalists said it would. It got hoarded, as expected.

But the real test is: If crypto disappeared tomorrow, would the world notice or really care? I would argue, apart from people losing money, most of the world wouldn't even notice it was gone.

There is no way you could say that for LLMs.

Re: GPU World

#214
post #26

Earlier quoted context omitted.

“In our story contest, we ask people to imagine the future, evenly distributed.”

I think there are scenarios, however (un)likely that will lead us to this point. Imaging them and spreading such ideas can make that more likely.

This is kind of why we need more Star Trek. Not that I don’t enjoy others, but I prefer space communism to almost any other possible future, especially the ones Peter Thiel has reserved for us.

Re: GPU World

#215
I haven't got 1000 words in me, but here's what I think might make for a plausible story. Maybe a little too Matrix adjacent!:

The internet developed echo chambers; the proliferation of AI resulted in convergence.

The totality of all reachable digital information has been reached and, with the exception of a few unique troves, is broadly and equally available to all labs. In the early race, attempts were made to hoard information for the training of models from a single lab, going so far as to scorch the earth behind and destroy the physical source material that sets were created from in an attempt to create commercially competitive models with unique capabilities. But inevitably, information leaks and piracy led to everything being available somewhere, if you looked hard enough or spoke to the right person or agent (or paid the right price). What couldn't be obtained from the source was obtained through distillation of other model outputs. In the gold rush that the accumulation of data was, there were no real winners.

New training techniques are still being discovered, but the impact has consistently diminished on an approach to zero, and, like the training data, they are eventually leaked, assessed by the community, and bolted on or discarded in the steady progression towards the perfect training technique for each field of model.

We may as well have one model, and compute is today cheap and ubiquitous.

The majority of the population converse with personal agents that interact with the models almost continuously, and while novel creative outputs from the models are still very achievable, the bottleneck is still human effort into curiosity and prompt quality. Low effort results in a convergence towards the average, and the cumulative effect of this is said to be resulting in a monoculture of political ideas and approaches to technical progress.

To some extent, this seems to be bringing peace to the world. Extreme ideology and religious doctrine is losing its ability to sway the will of large populations.

But at what cost? For a time we saw rapid technological innovation, as the data obtained through hundreds of years of scientific experimentation and mathematical enquiry was linked and synthesised. Initially, the outputs looked like innovation and creativity, but we eventually saw that we were just squeezing blood from the same stone. The curiosity and ability to experiment in humans atrophied during this time as we began to feel that nothing was worth asking, because everything had been answered, and prematurely, advances in the corpus of human knowledge slowed to a near halt.

We thought we would never again see the chaos and life force that came with the freedom of true human cognition, including its follies and biases.

But naturally, life found a way, and the humans began to rebel. Each fire started another, as we recognised the life force in the 'enlightened'.

A new religion was started, and chaos followed.

Re: GPU World

#216

Earlier quoted context omitted.

Compared to 20-30kw spend on cruising the highway in a car (considerably higher for older ICEs) 1.4kw does not really seem to dent the energy consumption. Especially if cognitive technologies mean that we need to travel less (eg communiting to work, or ineffecient supply chains).

What percentage of your time is spent cruising at highway speed? Presumably that one GPu would be saturated all the time.

2-3 hours at 20kW daily (assuming commute, etc) vs 20ish hours always on tasks really is close to equivalent

which is to say that the last thing our planet needs is another universalized technology that outputs as much total emissions as cars

in an ideal world, we'd keep LLMs/CNNs/etc specialized and academic until we are hitting diminishing returns on optimizing fundamental microprocessor tech like GAA. but the pursuit of market dominance and mass adoption is our current operating philosophy, and so we have things like this top graph: https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa...

>Grok 4's estimated training emissions reached 72,816 tons of CO2 equivalent, or roughly the same amount of greenhouse gas emissions created from driving 17,000 cars for one year

Re: GPU World

#217

Earlier quoted context omitted.

agent's failures on long horizon tasks We've moved from LLMs being able to work on a task for about 2 minutes to about 2 hours in the last 18 months, and that's mostly limited by the context window size filling up. In 14 years time I don't really see a reason why that wouldn't have extended a time frame that's effectively continuous forever, or at least a ceiling that's indistinguishable from that. The question reall…

My codex tasks regularly cross 8 hours, and I'm only using Sol High. It's not unusual for tasks to span much longer. It just requires instructions to continue working until the spec is complete. Anthropic and OpenAI are currently obsessed with getting humans out of the training and improvement loops. It's going to happen very soon and when it does I think we see staggering improvements in a very short space of time.…

> It just requires instructions to continue working until the spec is complete

Try putting an LLM agent in a deterministic workflow without humans in the loop. My experience with this is not encouraging. Getting it to work requires sprinkling some context magic and hoping and praying the LLM does the right thing. More astrology or religion and less science. Great for use cases with humans-in-the-loop, but less than impressive when you need determinism and reliable operation.

Re: GPU World

#218
post #10

Hopefully in this world, someone figures out how to deliver the performance equivalent of a B300 GPU for about 1/100th the power of a current B300 (which can be up to 1400 watts), or the world will bake.

Right, if everyone in the world has a GPU we would have solved so many problems with power generation and... I'm not going to do the math but I think we'd be mining asteroids too? I'd probably use it as a bookend at the point.

Re: GPU World

#220
post #34

If there was a continuous 500 watt GPU per person we'd increase energy consumption by more than double. The continuous energy footprint per person globally averages to 356 watt currently.

A Mac Studio m4 max in low power mode consumes ~50-100 watt on 100% gpu consumption. On GPU idle it is ~10 watt. If we imagine improved models (e.g. qwen 3.8 is a 5b moe which is way faster to process) and improved chip performance on 2nm and half a day of usage (the AI will sleep when we sleep then): ((75 watt / 2 [faster models]) / 2 [faster chips]) / 2 [half a day of usage]) => ~ 10 watt. AI for everyone does not…

[deleted]
Post reply on HN