Live data from Hacker News

The RAM shortage comes for us all

jeffgeerling.com

251–260 of 416 posts

Re: The RAM shortage comes for us all

#251

> And those companies all realized they can make billions more dollars making RAM just for AI datacenter products, and neglect the rest of the market. I wouldn't ascribe that much intent. More simply, datacenter builders have bought up the entire supply (and likely future production for some time), hence the supply shortfall. This is a very simple supply-and-demand situation, nothing nefarious about it.

That makes it sound like they are powerless, which is not the case. They don’t have to have their capacity fully bought out, they could choose to keep a proportion of capacity for maintaining the existing PC market, which they would do if they thought it would benefit them in the long term. They’re not doing that, because it benefits them not to.

$20B, 5 years and you can have your own DDR5 fab to print money with.

jokes aside, if the AI demand actually materializes, somebody will look at the above calculation and say 'we're doing it in 12 months' with a completely straight face - incumbents' margin will be the upstart's opportunity.

Re: The RAM shortage comes for us all

#252

Red chip supply problems in your factory are usually caused by insufficient plastic bars, which is usually caused by oil production backing up because you're not consuming your heavy oil and/or petroleum fast enough. Crack heavy oil to light, and turn excess petroleum into solid fuel. As a further refinement, you can put these latter conversions behind pumps, and use the circuit network to only turn the pumps on when…

Yup, and stockpiling solid fuel is not a waste because you need it for rocket fuel later on. Just add more chests.

Re: The RAM shortage comes for us all

#253
post #202

Earlier quoted context omitted.

What would you expect yield to be?

With no prior experience? 0%. Those machines are not just like printers :-)

Especially when the plan is to just run them in a random rented commercial warehouse.

I drive by a large fab most days of the week. A few breweries I like are down the street from a few small boutique fabs. I got to play with some experimental fab equipment in college. These aren't just some quickly thrown together spaces in any random warehouse.

And it's also ignoring the water manufacturing process, and having the right supply chain to receive and handle these ultra clean discs without introducing lots of gunk into your space.

Re: The RAM shortage comes for us all

#255
post #145
post #13

I think the OpenAI deal to lock wafers was a wonderful coup. OpenAI is more and more losing ground against the regularity[0] of the improvements coming from Anthropic, Google and even the open weights models. By creating a chock point at the hardware level, OpenAI can prevent the competition from increasing their reach because of the lack of hardware. [0]: For me this is really an important part of working with Claud…

Please explain to me like I am five: Why does OpenAI need so much RAM? 2024 production was (according to openai/chatgpt) 120 billion gigabytes. With 8 billion humans that's about 15 GB per person.

What they need is not so much memory but memory bandwidth.

For training, their models have a certain number of memory needed to store the parameters, and this memory is touched for every example of every iteration. Big models have 10^12 (>1T )parameters, and with typical values of 10^3 examples per batch, and 10^6 number of iteration. They need ~10^21 memory accesses per run. And they want to do multiple runs.

DDR5 RAM bandwidth is 100G/s = 10^11, Graphics RAM (HBM) is 1T/s = 10^12. By buying the wafer they get to choose which types of memory they get.

10^21 / 10^12 = 10^9s = 30 years of memory access (just to update the model weights), you need to also add a factor 10^1-10^3 to account for the memory access needed for the model computation)

But the good news is that it parallelize extremely well. If you parallelize you 1T parameters, 10^3 times, your run time is brought down to 10^6 s = 12 days. But you need 10^3 *10^12 = 10^15 Bytes of RAM by run for weight update and 10^18 for computation (your 120 billions gigabytes is 10^20, so not so far off).

Are all these memory access technically required : No if you use other algorithms, but more compute and memory is better if money is not a problem.

Is it strategically good to deprive your concurrents from access to memory : Very short-sighted yes.

It's a textbook cornering of the computing market to prevent the emergence of local models, because customers won't be able to buy the minimal RAM necessary to run the models locally even just the inferencing part (not the training). Basically a war on people where little Timmy won't be able to get a RAM stick to play computer games at Xmas.

Re: The RAM shortage comes for us all

#256
post #61

Every shortage is followed by a glut. Wait and see for RAM prices to go way down. This will happen because RAM makers are racing to produce units to reap profits from the higher price. That overproduction will cause prices to crash.

They aren't overproducing consumer modules, they're actively cutting production of those. They're producing datacenter/AI specific form factors that won't be compatible with consumer hardware.

somebody will step up to pick up the free money if this continues.

Re: The RAM shortage comes for us all

#257
post #241

Earlier quoted context omitted.

With no prior experience? 0%. Those machines are not just like printers :-)

We'll have to gain some experience then :)

Sure - once you have dozens of engineers and 5 years under your belt you'll be good to go!

This will get you started: https://youtu.be/B2482h_TNwg

Keep in mind that every wafer makes multiple trips around the fab, and on each trip it visits multiple machines. Broadly, one trip lays down one layer, and you may need 80-100 layers (although I guess DRAM will be fewer). Each layer must be aligned to nanometer precision with previous layers, otherwise the wafer is junk.

Then as others have said, once you finish the wafer, you still need to slice it, test the dies, and then package them.

Plus all the other stuff....

You'll need billions in investment, not millions - good luck!

Re: The RAM shortage comes for us all

#258

Perhaps we'll have to start optimizing software for performance and RAM usage again. I look at MS Teams currently using 1.5GB of RAM doing nothing.

RIP electron apps and PWAs. Need to go native, as chromium based stuff is so memory hungry. PWAs on Safari use way less memory, but PWA support in Safari is not great.

Re: The RAM shortage comes for us all

#259
post #124

This reminds me of the recent LaurieWired video presenting a hypothetical of, "what if we stopped making CPUs": https://www.youtube.com/watch?v=L2OJFqs8bUk Spoiler, but the answer is basically that old hardware rules the day because it lasts longer and is more reliable of timespans of decades. DDR5 32GB is currently going for ~$330 on Amazon DDR4 32GB is currently going for ~$130 on Amazon DDR3 32GB is currently goin…

Yes, DDR3 is the lowest CAS latency and lasts ALOT longer. Just like SSDs from 2010 have 100.000 writes per bit instead of below 10.000. CPUs might even follow the same durability pattern but that remains to be seen. Keep your old machines alive and backed up!

> 100.000 writes per bit

per cell*

Also, that SSD example is wildly untrue. Especially with the context of available capacity at the time. You CAN get modern SSD's with mind boggling write endurance per cell, AND has multides more cells, resulting in vastly more durable media than what was available pre 2015. The one caveat there to modern stuff being better than older stuff is Optane (the enterprise stuff like the 905P or P5800X, not that memory and SSD combo shitshow that Intel was shoveling out the consumer door). We still haven't reached parity with the 3DXpoint stuff, and it's a damn shame Intel hurt itself in it's confusion and cancelled that, because boy would they and Micron be printing money hand over fist right now if they were still making them. Still, Point being: Not everything is a TLC/QLC 0.3DWPD disposable drive like has become standard in the consumer space. If you want write endurance, capacity, and/or performance, you have more and better options today than ever before (Optane/3DXPoint excepted).

Regarding CPU's, they still follow that durability pattern if you unfuck what Intel and AMD are doing with boosting behavior and limit them to perform with the margins that they used to "back in the day". This is more of a problem on the consumer side (Core/Ryzen) than the enterprise side (Epyc/Xeon). It's also part of why the OC market is dying (save for maybe the XOC market that is having fun with LN2), those CPU's (especially consumer ones) come from the factory with much less margin for pushing things, because they're already close to their limit without exceedingly robust cooling.

I have no idea what the relative durability of RAM is tbh, it's been pretty bulletproof in my experience over the years, or at least bulletproof enough for my usecases that I haven't really noticed a difference. Notable exception is what I see in GPU's, but that is largely heat-death related and often a result of poor QA by the AIB that made it (eg, thermal pads not making contact with the GDDR modules).

Re: The RAM shortage comes for us all

#260
post #246
post #216

Earlier quoted context omitted.

You're just talking about a lithography machine. Patterning is one step out of thousands in a modern process (albeit an important one). There's plenty more stuff needed for a production line, this isn't a 3D printer but for chips. And that's just for the FEOL stuff, then you still need to do BEOL :). And packaging. And testing (accelerated/environmental, too). And failure analysis. And... Also, you know, there's a wh…

> 3D printer but for chips how about a farm of electron microscopes? these should work

Canon has been working on an alternative to EUV lithography called nanoimprint lithography. It would be a bit closer to the idea of having an inkjet printer make the masks to etch the wafers. It hasn't been proven in scale and there's a lot of thinking this won't really be useful, but it's neat to see and maybe the detractors are wrong.

https://global.canon/en/technology/nil-2023.html

https://newsletter.semianalysis.com/p/nanoimprint-lithograph...

They'll still probably require a good bit of operator and designer knowledge to work around whatever rough edges exist in the technology to keep yields high, assuming it works. It's still not a "plug it in, feed it blank wafers, press PRINT, and out comes finished chips!" kind of machine some here seem to think exist.

Post reply on HN