I have had the same question lingering, so I guess there are many more people like me and you benefiting from this thread!
Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
211–220 of 379 posts
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#212Earlier quoted context omitted.
Even is the AI bubble does not pops, your prediction about those servers being available on ebay in 10 years will likely be true, because some datacenters will simply upgrade their hardware and resell their old ones to third parties.
Would anybody buy the hardware though? Sure, datacenters will get rid of the hardware - but only because it's no longer commercially profitable run them, presumably because compute demands have eclipsed their abilities. It's kind of like buying a used GeForce 980Ti in 2025. Would anyone buy them and run them besides out of nostalgia or curiosity? Just the power draw makes them uneconomical to run. Much more likely ev…
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#213Earlier quoted context omitted.
Would anybody buy the hardware though? Sure, datacenters will get rid of the hardware - but only because it's no longer commercially profitable run them, presumably because compute demands have eclipsed their abilities. It's kind of like buying a used GeForce 980Ti in 2025. Would anyone buy them and run them besides out of nostalgia or curiosity? Just the power draw makes them uneconomical to run. Much more likely ev…
> Sure, datacenters will get rid of the hardware - but only because it's no longer commercially profitable run them, presumably because compute demands have eclipsed their abilities. I think the existence of a pretty large secondary market for enterprise servers and such kind of shows that this won't be the case. Sure, if you're AWS and what you're selling _is_ raw compute, then couple generation old hardware may not…
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#214Earlier quoted context omitted.
Yes. Google is probably gonna win the LLM game tbh. They had a massive head start with TPUs which are very energy efficient compared to Nvidia Cards.
Yeah honestly. They could just try selling solutions and SLAs combining their TPU hardware with on-prem SOTA models and practically dominate enterprise. From what I understand, that's GCP's gameplay too for most regulated enterprise clients.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#215Earlier quoted context omitted.
Yes. Google is probably gonna win the LLM game tbh. They had a massive head start with TPUs which are very energy efficient compared to Nvidia Cards.
But they’re ASICs so any big architecture changes will be painful for them right?
So it isn't like Google designed a TPU for a specific model or architecture. They're pretty general purpose in a narrow field (oxymoron, but you get the point).
The set of operations Google designed into a TPU is very similar to what nvidia did, and it's about as broadly capable. But Google owns the IP and doesn't pay the premium and gets to design for their own specific needs.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#216I work at Google on these systems everyday (caveat this is my own words not my employers)). So I simultaneously can tell you that its smart people really thinking about every facet of the problem, and I can't tell you much more than that. However I can share this written by my colleagues! You'll find great explanations about accelerator architectures and the considerations made to make things fast. https://jax-ml.git…
> So I simultaneously can tell you that its smart people really thinking about every facet of the problem, and I can't tell you much more than that. "we do 1970s mainframe style timesharing" there, that was easy
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#217Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#218An H100 is a $20k USD card and has 80GB of vRAM. Imagine a 2U rack server with $100k of these cards in it. Now imagine an entire rack of these things, plus all the other components (CPUs, RAM, passive cooling or water cooling) and you're talking $1 million per rack, not including the costs to run them or the engineers needed to maintain them. Even the "cheaper" I don't think people realize the size of these compute u…
After years of “AI is a bubble, and will pop when everyone realizes they’re useless plagiarism parrots” it’s nice to move to the “AI is a bubble, and will pop when it becomes completely open and democratized” phase
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#219Earlier quoted context omitted.
Googles bread and butter is advertising, so they have a huge interest in keeping things in house. Data is more valuable to them than money from hardware sales. Even then, I think that their primary use case is going to be consumer grade good AI on phones. I dunno why Gemma QAT model fly so low on the radar, but you can basically get full scale Llamma 3 like performance from a single 3090 now, at home.
https://www.cnbc.com/2025/04/09/google-will-let-companies-ru... Google has already started the process of letting companies self-host Gemini, even on NVidia Blackwell GPUs. Although imho, they really should bundle it with their TPUs as a turnkey solution for those clients who haven't invested in large scale infra like DCs yet.
And also, Google's track record with hardware.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#220One clever ingredient in OpenAI's secret sauce is billions of dollars of losses. About $5 billion dollars lost in 2024. https://www.cnbc.com/2024/09/27/openai-sees-5-billion-loss-t...
One can serve a lot if models if allowed to burn through over a billion dollars with no profit requirement. Classic, VC-style, growth-focused capitalism with an unusual, business structure.