Live data from Hacker News

Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

tomshardware.com

161–170 of 331 posts

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#161

Key paragraph: > However, a small handful of models such as Alibaba’s Qwen and DeepSeek are most popular for inference, with most other models only sporadically called upon. This leads to resource inefficiency, with 17.7 per cent of GPUs allocated to serve only 1.35 per cent of requests in Alibaba Cloud’s marketplace, the researchers found.

these other models are likely much smaller

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#162

Earlier quoted context omitted.

>the same time the US is breaking up its own alliances This isn't happening. The US is driving a harder bargain with our allies. No one serious thinks anyone is walking away from alliances with the US.

Why the framing of alliances like it’s a boolean? The question of “can we trust the American government” is now being asked more often. Existing alliances and new potential alliances face that question, whether or not you personally believe that they should trust America. Even if no concrete actions are being performed with asking that question, the fact that question is even being asked is a major drop from where we…

You're right, alliances are not boolean.

From the US perspective, we have been asking ourselves "can we trust Europe's military capacity" for a very long time and the answer (prior to 2025) was: NO.

With Trump on one side and Russia on the other, it seems like the answer has shifted to: MAYBE.

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#163

Earlier quoted context omitted.

Re: Western. A similar thing plays out when the term "international community" is used in news. It refers to the US and its major allies which means US, Canada, Western Europe, Japan, Australia and New Zealand more or less.

> A similar thing plays out when the term "international community" is used in news. It refers to the US and its major allies which means US, Canada, Western Europe, Japan, Australia and New Zealand more or less. Wait, really? I thought "international community" meant all countries.

There was a particularly memorable use of this sense some time ago, when the UK representative to the UN explained that they abstained from a vote in the General Council that passed with something like 200+ members voting for it because "the international community is still divided on the topic".

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#164

Earlier quoted context omitted.

Re: Western. A similar thing plays out when the term "international community" is used in news. It refers to the US and its major allies which means US, Canada, Western Europe, Japan, Australia and New Zealand more or less.

Yes community refers to whose who participate in community. How is this hard to understand? Broadly speaking coast de ivory and the like is not a participant in the international community.

China, Russia, India, Pakistan, Iran, Saudi Arabia, and many many other countries that are very active members of the international community are not counted among members of THE "international community". Hell, much of Europe isn't either, including some of the former colonial empires, on some topics.

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#165

Alibaba Cloud claims to reduce Nvidia GPU used for serving unpopular models by 82% (emphasis mine) > 17.7 per cent of GPUs allocated to serve only 1.35 per cent of requests in Alibaba Cloud’s marketplace, the researchers found Instead of 1192 GPUs they now use 213 for serving those requests.

In the past, software and computer engineers would tackle problems head-on, designing algorithms and finding creative solutions. thanks to the US restrictions on semiconductor industry (Chinese), Chinese engineers are being forced to innovate and find their own ways to overcome challenges like the old school engineers (What Silicon Valley used to be)

If you're one who sees progress as an end goal unto itself, what you describe is a good thing. When one party is attempting novel solutions to outcompete the competition we will be faster to whatever the next change is.

That said, I'm not sure what the US policies specifically have to do with this. Countries are always in competition with one another, and if one industry or technology is considered a national security threat they will guard it.

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#166
post #41

Earlier quoted context omitted.

> If for purely political reasons, China may never be able to export these chips to most of the world - which limits their scale - which makes it harder to make them cost effective compared to Western chips. Note that this happens at the same time the US is breaking up its own alliances, so as of this writing, there's no such thing as certainty about politics.

>the same time the US is breaking up its own alliances This isn't happening. The US is driving a harder bargain with our allies. No one serious thinks anyone is walking away from alliances with the US.

I observe serious financial commitments towards walking away from US tech:

The EU is pumping money into what they call "digital sovereignty" left and right. Germany just cancelled their Microsoft subscriptions and replaced them with self-funded Open Source for Schleswig-Holstein, which is roughly 5% of all government employees. That's one hell of a trial run. Germany's "OpenDesk" and France’s "La Suite numérique" even made into the new "Franco-German Economic Agenda 2025", which self-describes as "bilateral coordination to full swing for a more sovereign Europe".

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#167

Alibaba Cloud claims to reduce Nvidia GPU used for serving unpopular models by 82% (emphasis mine) > 17.7 per cent of GPUs allocated to serve only 1.35 per cent of requests in Alibaba Cloud’s marketplace, the researchers found Instead of 1192 GPUs they now use 213 for serving those requests.

I’m slightly confuse as to how all this works. Do the GPUs just sit there with the models on them when the models are not in use? I guess I’d assumed this sort of thing would be allocated dynamically. Of course, there’s a benefit to minimizing the number of times you load a model. But surely if a GPU+model is idle for more than a couple minutes it could be freed? (I’m not an AI guy, though—actually I’m used to asking…

> I guess I’d assumed this sort of thing would be allocated dynamically

At the scale of a hyperscaler I think Alibaba is the one that would be doing that. AWS, Azure and I assume Alibaba do lease/rent data centers, but someone has to own the servers / GPU racks. I know there are specialized companies like nscale (and more further down the chain) in the mix, but I always assumed they only lease out fixed capacity.

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#168

Earlier quoted context omitted.

False. Westernism is broadly an extension of the academic notion of classicism, starting in Egypt and then Greece Rome and into Europe and the Americas.

It's not an academic notion (at least not strictly), since virtually everyone uses it routinely.

Oddly since I got many downvotes with this statement, it's clear the average hacker news reader knows very little about world history or common knowledge

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#169
post #63

Earlier quoted context omitted.

Not really, Figure 1(a) of the paper says that the 17.7% are relative to a total of 30k GPUs (i.e. 5310 GPUs for handling those 1.35% of requests) and the reduction is measured in a smaller beta deployment with only 47 different models (vs. the 733 "cold" models overall.) Naïve extrapolation by model count suggests they would need 3321 GPUs to serve all cold models, a 37.5% reduction to before. (Or 6.6% reduction of…

Really: "A paper presented at SOSP 2025 details how token-level scheduling helped one GPU serve multiple LLMs, reducing demand from 1,192 to 213 H20s." Which, if you scale it, matches the GPs statement.

From the SCMP article you might get the impression that the various figures all refer to the same GPU cluster, but in the paper itself it's very clear that this is not the case, i.e. the 213 GPUs in the smaller cluster are not serving 1.35% of the requests in the larger cluster. Then if you want to scale it, you have a choice of different numbers you could scale, and each would get different results. Since they're constrained by the limited number of different models a single GPU can serve, I think scaling by the number of models is the most realistic option.

Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system

#170
post #72

Earlier quoted context omitted.

I haven't read the paper yet, but it's here: https://ennanzhai.github.io/pub/sosp25-aegaeon.pdf aka https://dl.acm.org/doi/10.1145/3731569.3764815 . So, definitely not state media, probably not lying on the fundamentals. Of course, still presumably viewed favorably by the CCP, I'd imagine.

Well if it is real, we will surely see out Claude Code limits go back up.

From the abstract, this seems to be a scheduling mechanism for datacenters that serve multiple models. I have no idea whether this applies to Anthropic.
Post reply on HN