Earlier quoted context omitted.
But at the same time they are fielding multiple new stealth aircraft and their jets and missiles outperformed western aircraft in the recent Pakistan India flare-up.
So you'd think they'd be able to build a commercial jet liner, no?
Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
111–120 of 331 posts
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#112Earlier quoted context omitted.
> The streets are flooded with cheap Chinese cars and I see more BYD than American cars This isn't surprising in any way, American "cars" (quotes because the vast majority of what American manufacturers pump out isn't cars, it's trucks) haven't been competitive in decades. The only globally competitive vehicles were developed in Europe by GM Europe (Opel, since sold to PSA now Stellantis) or Ford Europe (which axed a…
The US couldn't compete anyway. The US is an advanced economy and will struggle to compete in non-advanced categories. It's like trying to level your MMORPG character to 100 by only farming in lvl 30-40 mob areas. It's really not worth it and mostly forced.
Take Renault for example, their Renault 5 and 4 EVs are good looking, not luxury but definitely premium, and the 5 sedan starts at 30k€; the 4 crossover starts at 29k€. This is before a 5k€ government subsidy. Their boring, fewer bells and whistles, low cost model, the Dacia Spring, starts at 17k€. The Renault 5 and 4 are made almost entirely in France, while the Dacia is made in Romania - a lower cost country, but still an EU member state.
The comparable in size and autonomy BYD Dolphin starts at 20k€. Both for cheapness and quality/design, Renault are competitive.
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#113better link https://www.tomshardware.com/tech-industry/semiconductors/al... paper https://dl.acm.org/doi/10.1145/3731569.3764815
Ok, we've changed the URL above (from https://www.scmp.com/business/article/3329450/alibaba-cloud-... ), and will put the link to the paper in the top text. Thanks!
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#114Earlier quoted context omitted.
But at the same time they are fielding multiple new stealth aircraft and their jets and missiles outperformed western aircraft in the recent Pakistan India flare-up.
So you'd think they'd be able to build a commercial jet liner, no?
See what I did there?
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#115Earlier quoted context omitted.
Tbh this whole situation reminds of how Japan excelled in making a lot more with a lot less after WW2, e.g., fuel-efficient engines, light cars, etc. these constraints were not present in the US (and to some extent in Europe), and resulted in US cars being completely not competitive in non-US markets.
The premature optimizer is never the innovator. Japan eventually stopped that role and their products improved greatly.
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#116Earlier quoted context omitted.
The US couldn't compete anyway. The US is an advanced economy and will struggle to compete in non-advanced categories. It's like trying to level your MMORPG character to 100 by only farming in lvl 30-40 mob areas. It's really not worth it and mostly forced.
Cars are relatively advanced and more importantly, seen as a status symbol by many. There is, IMO, plenty of space for not-the-cheapest, quality cars. Take Renault for example, their Renault 5 and 4 EVs are good looking, not luxury but definitely premium, and the 5 sedan starts at 30k€; the 4 crossover starts at 29k€. This is before a 5k€ government subsidy. Their boring, fewer bells and whistles, low cost model, the…
They really nailed the modern-with-subtle-calls-to-retro look.
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#117Earlier quoted context omitted.
Re: Western. A similar thing plays out when the term "international community" is used in news. It refers to the US and its major allies which means US, Canada, Western Europe, Japan, Australia and New Zealand more or less.
> A similar thing plays out when the term "international community" is used in news. It refers to the US and its major allies which means US, Canada, Western Europe, Japan, Australia and New Zealand more or less. Wait, really? I thought "international community" meant all countries.
Sometimes it's used in the expected way, but (more?) often, "international community" euphemistically refers to whomever is currently one of, or an ally of the above mentioned countries.
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#118Earlier quoted context omitted.
Cars are relatively advanced and more importantly, seen as a status symbol by many. There is, IMO, plenty of space for not-the-cheapest, quality cars. Take Renault for example, their Renault 5 and 4 EVs are good looking, not luxury but definitely premium, and the 5 sedan starts at 30k€; the 4 crossover starts at 29k€. This is before a 5k€ government subsidy. Their boring, fewer bells and whistles, low cost model, the…
The new 5 is one of the first cars I've really liked in a while. They really nailed the modern-with-subtle-calls-to-retro look.
Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#119Re: Alibaba Cloud says it cut Nvidia AI GPU use by 82% with new pooling system
#120Alibaba Cloud claims to reduce Nvidia GPU used for serving unpopular models by 82% (emphasis mine) > 17.7 per cent of GPUs allocated to serve only 1.35 per cent of requests in Alibaba Cloud’s marketplace, the researchers found Instead of 1192 GPUs they now use 213 for serving those requests.
Not really, Figure 1(a) of the paper says that the 17.7% are relative to a total of 30k GPUs (i.e. 5310 GPUs for handling those 1.35% of requests) and the reduction is measured in a smaller beta deployment with only 47 different models (vs. the 733 "cold" models overall.) Naïve extrapolation by model count suggests they would need 3321 GPUs to serve all cold models, a 37.5% reduction to before. (Or 6.6% reduction of…
"A paper presented at SOSP 2025 details how token-level scheduling helped one GPU serve multiple LLMs, reducing demand from 1,192 to 213 H20s."
Which, if you scale it, matches the GPs statement.