Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
291–300 of 379 posts
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#292Earlier quoted context omitted.
The only one who can stop Google is Google. They’ll definitely have the best model, but there is a chance they will f*up the product / integration into their products.
It would take talent for them to mess up hosting businesses who want to use their TPUs on GCP. But then again even there, their reputation for abandoning products, lack of customer service, condescension when it came to large enterprises’ “legacy tech” lets Microsoft who is king of hand holding big enterprise and even AWS run rough shod over them. When I was at AWS ProServe, we didn’t even bother coming up with talki…
What are the chances of abandoning TPU-related projects where the company literally invested billions in infrastructure? Zero.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#293Earlier quoted context omitted.
Yes. Google is probably gonna win the LLM game tbh. They had a massive head start with TPUs which are very energy efficient compared to Nvidia Cards.
Google will win the LLM game if the LLM game is about compute , which is the common wisdom and maybe true, but not foreordained by God. There's an argument that if compute was the dominant term that Google would never have been anything but leading by a lot. Personally right now I see one clear leader and one group going 0-99 like a five sigma cosmic ray: Anthropic and the PRC. But this is because I believe/know that…
All the LLM vendors are going to have to cope with the fact that they're lighting money on fire, and Google have the paying customers (advertisers) and with the user-specific context they get from their LLM products, one of the juciest and most targetable ad audiences of all time.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#294I work at Google on these systems everyday (caveat this is my own words not my employers)). So I simultaneously can tell you that its smart people really thinking about every facet of the problem, and I can't tell you much more than that. However I can share this written by my colleagues! You'll find great explanations about accelerator architectures and the considerations made to make things fast. https://jax-ml.git…
If people at google are so smart why can't google.com get a 100% lighthouse score?
I don't like how the grand parent mystifies this. This problem is just normal engineering. Any good engineer could learn how to do it.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#295Earlier quoted context omitted.
If people at google are so smart why can't google.com get a 100% lighthouse score?
Because most smart people are not generalists. My first boss was really smart and managed to found a university institute in computer science. The 3 other professors he hired were, ahem, strange choices. We 28 year old assistents could only shake our heads. After fighting a couple of years with his own hires the founder left in frustration to found another institution. One of my colleagues was only 25, really smart i…
The real answer is likely internal company politics and priorities. Google certainly has people with the technical skills to solve it but do they care and if they care can they allocate those skilled people to the task?
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#296Earlier quoted context omitted.
This is the real answer, I don't know what people above are even discussing when batching is the biggest reduction in costs. If it costs say $50k to serve one request, with batching is also costs $50k to serve 100 at the same time with minimal performance loss, I don't know what the real number of users is before you need to buy new hardware, but I know it's in the hundreds so going from $50000 to $500 in effective c…
Thanks for the helpful reply! As I wasn't able to fully understand it still, I pasted your reply in chatgpt and asked it some follow up questions and here is what i understand from my interaction: - Big models like GPT-4 are split across many GPUs (sharding). - Each GPU holds some layers in VRAM. - To process a request, weights for a layer must be loaded from VRAM into the GPU's tiny on-chip cache before doing the ma…
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#297Earlier quoted context omitted.
Thanks for the helpful reply! As I wasn't able to fully understand it still, I pasted your reply in chatgpt and asked it some follow up questions and here is what i understand from my interaction: - Big models like GPT-4 are split across many GPUs (sharding). - Each GPU holds some layers in VRAM. - To process a request, weights for a layer must be loaded from VRAM into the GPU's tiny on-chip cache before doing the ma…
This seems a bit complicated to me. They don't serve very many models. My assumption is they just dedicate GPUs to specific models, so the model is always in VRAM. No loading per request - it takes a while to load a model in anyway. The limiting factor compared to local is dedicated VRAM - if you dedicate 80GB of VRAM locally 24 hours/day so response times are fast, you're wasting most of the time when you're not que…
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#298Earlier quoted context omitted.
This assumes that the value obtained by customers is high enough to cover any possible actual cost. Many current AI uses are low value things or one time things (for example CV generation, which is killing online hiring).
Many current AI uses are low value things or one time things (for example CV generation, which is killing online hiring). We are talking about Pro subs who have high usage.
At the end of the day, until at least one of the big providers gives us balance sheet numbers, we don't know where they stand. My current bet is that they're losing money whichever way you dice it.
The hope being as usual that costs go down and the market share gained makes up for it. At which point I wouldn't be shocked by pro licenses running into the several hundred bucks per month.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#299Earlier quoted context omitted.
Because most smart people are not generalists. My first boss was really smart and managed to found a university institute in computer science. The 3 other professors he hired were, ahem, strange choices. We 28 year old assistents could only shake our heads. After fighting a couple of years with his own hires the founder left in frustration to found another institution. One of my colleagues was only 25, really smart i…
I have met those supersmart specialists but in my experience there are also a lot of smart people who are more generalists. The real answer is likely internal company politics and priorities. Google certainly has people with the technical skills to solve it but do they care and if they care can they allocate those skilled people to the task?
It’s quite intimidating how fast they can break down difficult concepts into first principles. I’ve witnessed this first hand and it’s beyond intimidating. Makes you wondering what you’re doing at this company… That being said, the caliber of folks I’m talking about is quite rare, like top 10% of top 1% teams at Google.
Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
#300Earlier quoted context omitted.
It would take talent for them to mess up hosting businesses who want to use their TPUs on GCP. But then again even there, their reputation for abandoning products, lack of customer service, condescension when it came to large enterprises’ “legacy tech” lets Microsoft who is king of hand holding big enterprise and even AWS run rough shod over them. When I was at AWS ProServe, we didn’t even bother coming up with talki…
> It would take talent for them to mess up hosting businesses who want to use their TPUs on GCP. > But then again even there, their reputation for abandoning products What are the chances of abandoning TPU-related projects where the company literally invested billions in infrastructure? Zero.
All things that Google is remarkably bad at.