Live data from Hacker News

Microsoft eyes $10B bet on ChatGPT

semafor.com

521–530 of 541 posts

Re: Microsoft eyes $10B bet on ChatGPT

#521
post #256

Earlier quoted context omitted.

The closest open source contender is BLOOM: https://huggingface.co/bigscience/bloom . It has an almost identical architecture to GPT-3 (hence, to ChatGPT), and in particular the same number of parameters (175B). It was also trained on a similar amount of data, as far as we can know. Still, it's not like you can just "download it and run it", even just to _load_ the model into memory you need ~400GB of memory, to run…

Getting a server with > 400GB of RAM and a heap of GPUs can be done for less than $6,000 - $10,000 if you're scrappy. Not cheap, but also not out of reach for individuals.

I don't think that figure is correct, you need a "good heap" of GPUs, not just anything... in particular, even just to run inference, you need at least 400 GB of GPU memory, not just RAM. You can't just plug a dozen "cheap" GPUs and call it a day, because if I remember correctly consumer GPUs have at most 32GB of RAM each. Hence you'd need at least 12 of those top-tier GPUs (which certainly don't come at $500 a piece). Probably more, because you can't trivially split weights across GPUs so perfectly (you probably have to put an integer number of layers on each GPU).

In practice these models are typically run using top-tier A100 GPUs, which apparently is the cheapest thing you can do at scale: https://forum.effectivealtruism.org/posts/foptmf8C25TzJuit6/.... It looks like you can get away with just $10/hour, but I'm not sure I believe it. In one hour you can roughly generate 6 million English words this way, that's quite cheap.

But if you want to own the full hardware, then it's quite more expensive. You need 8 of those A100 GPUs, which come at $32k a piece, so you're in the ballpark of > $300k to build the server you need. Then there's of course running costs, these GPUs burn 250W a piece, plus the rest of the server we're at about 3kW power. That's not much, maybe $0.50/hr, plus maybe another $1/hr to cool the room it's in, depending on where it is (and the season, I guess in winter a fan might suffice, it's about as powerful as a couple small electric heaters). So with an upfront expense of > $300k, you're maybe down from $10/hr to $1.5/hr, saving something like $8.5/hr, which is $6k / month (minus the rent of whatever place you put the server in).

All in all, it's definitely feasible for a small start up as well, but not very much for an individual.

Re: Microsoft eyes $10B bet on ChatGPT

#522

Earlier quoted context omitted.

That's cheap for an AI model training at big tech. Your average ads model engineer at those companies likely uses more compute each quarter. It mostly suggests that we're in another AI hype bubble. MS and other big tech companies can easily replicate OpenAI results and do same type of research.

The number of companies ready to spend $40M/year of compute per "average ads model engineer" is exactly 0.

I know of 1 at least...

'compute' costs are tricky to convert to dollar values - everyone does it at 'what would it cost to rent this from AWS', but the reality is, people like AWS are prepared to give low priority access to unsold compute for almost any project.

Re: Microsoft eyes $10B bet on ChatGPT

#524
post #511

Earlier quoted context omitted.

No, the GNU licences place copyleft obligations on distribution/conveyance. But they allow you to run the programs for any purpose, without field of endeavour restrictions, or moral police. You don't even need to accept a licence to run a GNU program.

Out of interest are there any copyleft style neural network licenses - eg that require fine tuned model weights are published? (And Affero GPL style in terms of servers and distribution meaning these days)

I am not an expert, but I'd just use the the normal LGPL/GPL/AGPL licenses for the models.

Re: Microsoft eyes $10B bet on ChatGPT

#525
post #395

Earlier quoted context omitted.

Compare https://opensource.org/osd This isn't a very minor point, as this was an explicit discussion and is also OSI's translation of Debian's translation of Richard Stallman's "freedom 0". That is, it's an important, and explicit, tradition/consensus in FOSS that users aren't restricted in the purposes for which they may use the software.

I don't have an opinion on this one way or another, but if the RAIL license concerns you, then perhaps you can take it up with the organization behind it? https://www.licenses.ai/

There's no point. RAIL are well aware their licenses are proprietary (see their FAQ), and they are happy with that.

Re: Microsoft eyes $10B bet on ChatGPT

#526

> Additionally, a structure with the ownership of OpenAI would be put in place. Microsoft would entail a 49 percent stake while other investors would take over the other 49 percent. The remaining 2 percent will reportedly go to OpenAI’s non-profit parent firm. When OpenAI was funded as a non-profit in 2015, it raised $1 billion from investors that included YC Research. [1] Sam Altman was also the former president of…

One thing that really bothers me is how much of a contradiction "OpenAI" is. There's virtually nothing open about it.

Disagreed. I'd say there's almost nothing open about it. GPT-1, GPT-2, Whisper, Point e, various papers etc. were all "open" contributions.

Re: Microsoft eyes $10B bet on ChatGPT

#527
post #521

Earlier quoted context omitted.

Getting a server with > 400GB of RAM and a heap of GPUs can be done for less than $6,000 - $10,000 if you're scrappy. Not cheap, but also not out of reach for individuals.

I don't think that figure is correct, you need a "good heap" of GPUs, not just anything... in particular, even just to run inference, you need at least 400 GB of GPU memory, not just RAM. You can't just plug a dozen "cheap" GPUs and call it a day, because if I remember correctly consumer GPUs have at most 32GB of RAM each. Hence you'd need at least 12 of those top-tier GPUs (which certainly don't come at $500 a piece…

Got it, thanks for the information! I hadn't known it was all VRAM for model serving.

Re: Microsoft eyes $10B bet on ChatGPT

#528

Earlier quoted context omitted.

Creating new life forms? Many people can barely raise their own children. Can you imagine if tech companies create advanced lifeforms? It's truly a disaster waiting to happen. It's rather rude to call it pointless FUD. AI messing up our world is real and I have thought it through at great length. Having the perfect friend is actually immensely scary. It will create a world where people interact mostly with their elec…

> Creating new life forms? Many people can barely raise their own children. Can you imagine if tech companies create advanced lifeforms? It's truly a disaster waiting to happen. I never said it would be a good thing per se. >It's rather rude to call it pointless FUD. AI messing up our world is real and I have thought it through at great length. The particular comment I was referring to was FUD about humans losing the…

[deleted]

Re: Microsoft eyes $10B bet on ChatGPT

#529

Earlier quoted context omitted.

> Creating new life forms? Many people can barely raise their own children. Can you imagine if tech companies create advanced lifeforms? It's truly a disaster waiting to happen. I never said it would be a good thing per se. >It's rather rude to call it pointless FUD. AI messing up our world is real and I have thought it through at great length. The particular comment I was referring to was FUD about humans losing the…

> You can't really think this a possible or reasonable undertaking. Luddism isn't going to save us. Yes, I do. I believe forsaking a lot of new technology and curbing technological innovation is an excellent way to further humanity. I blog about it and write a newsletter and talk to anyone who will listen.

>Yes, I do. I believe forsaking a lot of new technology and curbing technological innovation is an excellent way to further humanity.

The problem is that the cat is out of the bag. We didn't pass laws saying you can't make chip foundries and now they are all over the world. Even a Nation State would have issues destroying all of them and now the reality is that even non-ideal architectures can be used to train models. So it's just not possible to stop progress. Even if you blew up every fab in the world some clever hackers are going to string together 100 Tesla Model S' or 20,000 smart fridges and train a Neural Net on them.

Blowing up fabs to stop progress is just not going to work not to mention the devastating effects it would have on the rest of the economy.

Re: Microsoft eyes $10B bet on ChatGPT

#530

Earlier quoted context omitted.

Interesting. What’s the connection?

They are darker skinned than the rest of the population and the word itself is like a generic name for one of them. I guess it is some form of mutation of the Turkish word mangal which is something like barbeque or charcoal

Thanks. Very interesting.
Post reply on HN