Live data from Hacker News

GPT-4 API General Availability

openai.com

151–160 of 562 posts

Re: GPT-4 API General Availability

#151

Earlier quoted context omitted.

I felt the same thing. The first version of GPT-4 I tried was crazy smart. Scary smart. Something happened afterwards…

The even more interesting part is that none of us got to try the internal version which was allegedly yet another step above that.

Oh, it's not too hard to see how the spend that Microsoft put into building the data centers where GPT-4 was trained attracted national security interest even before it went public. The fact that they were even allowed to release it publicly is likely due to its strategic deterrence effect and that they believed the released version was already a dumbed-down version.

The fact that rumors about GPT-5 were quickly suppressed and the models were dumbed down even more cannot be entirely explained by excessive demand. I think it's more likely that GPT-3.5 and GPT-4 demonstrated unexpected capabilities in the hands of the public leading to a pull back. Moreover, Sam Altman's behaviors changed dramatically between the initial release and a few weeks afterward -- the extreme optimism of a CEO followed by a more subdued, even cowed, demeanor despite strong enthusiasm from end-users.

OpenAI cannot do anything without Microsoft's data center resources, and Microsoft is a critical defense contractor.

Anyway, personally, I'm with the crowd that thinks we're about to see a Cambrian explosion of domain-specific expert AIs. I suspect that OpenAI/Microsoft/Gov is still trying to figure out how much to nerf the capability of GPT-3.5 to tutor smaller models (see "Textbooks are all you need") and that's why the API is trash.

Re: GPT-4 API General Availability

#152

Earlier quoted context omitted.

There's a big thread on ChatGPT getting dumber over on the ChatGPT subreddit, where someone suggests this is from model quantization: https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq... I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI…

$10-$20 per query? Can I get some sourcing on that? That's astronomically expensive.

I would presume that number includes the amortized training cost.

Re: GPT-4 API General Availability

#153

Earlier quoted context omitted.

Quick note: your domain doesn't appear to have an A record. I was hoping to follow the link in your profile and see if you have anything interesting written about LLMs.

Thanks! The website is no longer active, just updated my bio.

I know you guys are busy literally building the future but could you consider adding a search field in ChatGPT so that users can search their previous chats?

Re: GPT-4 API General Availability

#154

Earlier quoted context omitted.

There's a big thread on ChatGPT getting dumber over on the ChatGPT subreddit, where someone suggests this is from model quantization: https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq... I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI…

$10-$20 per query? Can I get some sourcing on that? That's astronomically expensive.

yeah this isnt close. Sam Altman is on record saying its single digit cents per query and then took a massively dilutive $10b investment from microsoft. Even if gpt4 is 8 models in a trenchcoat they wouldnt raise it on themselves by 4 orders of magnitude like that

Re: GPT-4 API General Availability

#155

I imagine the API quality isnt nerfed on a given day like ChatGPT can be. There was no question something happened in January with ChatGPT, weirdly would refuse to answer questions that were harmless but difficult(Give me a daily schedule of a stoic hedonist) Every once in a while, I see redditors complain of it being nerfed. Sometimes I go back to gpt3.5 and am mind boggled how much worse it is. Makes me wonder if t…

Instead of the model changing, it’s equally likely that this is a cognitive illusion. A new model is initially mind-blowing and enjoys a halo effect. Over time, this fades and we become frustrated with the limitations that were there all along.

But given the rumored architecture (MoE) it would make complete sense for them to dynamically scale down the number of models used in the mixture during periods of peak load.

Re: GPT-4 API General Availability

#156

Earlier quoted context omitted.

There's a big thread on ChatGPT getting dumber over on the ChatGPT subreddit, where someone suggests this is from model quantization: https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq... I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI…

$10-$20 per query? Can I get some sourcing on that? That's astronomically expensive.

There is absolutely no way. You can run a halfway decent open source model on a gpu for literally pennies in amortized hardware / energy cost.

Re: GPT-4 API General Availability

#157
post #89

Practical report: the OpenAI API is a bad joke. If you think you can build a production app against it, think again. I've been trying to use it for the past 6 weeks or so. If you use tiny prompts, you'll generally be fine (that's why you always get people commenting that it works for them), but just try to get closer to the limits, especially with GPT-4. The API will make you wait up to 10 minutes, and then time out.…

FWIW we have a live product for all users against gpt-3.5-turbo and it's largely fine: https://www.honeycomb.io/blog/improving-llms-production-obse...

In our own tracking, the P99 isn't exactly great, but this is groundbreaking tech we're dealing with here, and our dissatisfaction with the high end of latency is well worth the value we get in our product: https://twitter.com/_cartermp/status/1674092825053655040/

Re: GPT-4 API General Availability

#158
post #89

Practical report: the OpenAI API is a bad joke. If you think you can build a production app against it, think again. I've been trying to use it for the past 6 weeks or so. If you use tiny prompts, you'll generally be fine (that's why you always get people commenting that it works for them), but just try to get closer to the limits, especially with GPT-4. The API will make you wait up to 10 minutes, and then time out.…

[flagged]

Can you please not post in the flamewar style? We're trying for something else here and you can make your substantive points without it.

https://news.ycombinator.com/newsguidelines.html

Re: GPT-4 API General Availability

#159

Earlier quoted context omitted.

There's a big thread on ChatGPT getting dumber over on the ChatGPT subreddit, where someone suggests this is from model quantization: https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq... I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI…

$10-$20 per query? Can I get some sourcing on that? That's astronomically expensive.

People theorize that queries are being run on multiple A100's, each with a $10k ASP.

If you assume an A100 lives at the cutting edge for 2 years, that's about a million minutes, or $0.01 per minute of amortized HW cost.

In the crazy scenarios, I've heard 10 A100s per query, so assuming that takes a minute, maybe $0.1 per query.

Add an order of magnitude on top of that for labor/networking/CPU/memory/power/utilization/general datacenter stuff, you get to maybe $1/query.

So probably not $10, but maybe if you amortize training, low to mid single digits dollars per query?

Re: GPT-4 API General Availability

#160
post #89

Practical report: the OpenAI API is a bad joke. If you think you can build a production app against it, think again. I've been trying to use it for the past 6 weeks or so. If you use tiny prompts, you'll generally be fine (that's why you always get people commenting that it works for them), but just try to get closer to the limits, especially with GPT-4. The API will make you wait up to 10 minutes, and then time out.…

There's a big thread on ChatGPT getting dumber over on the ChatGPT subreddit, where someone suggests this is from model quantization: https://www.reddit.com/r/ChatGPT/comments/14ruui2/comment/jq... I've heard LLMs described as "setting money on fire" from people that work in the actually-running-these-things-in-prod industry. Ballpark numbers of $10-20/query in hardware costs. Right now Microsoft (through its OpenAI…

Note that /r/ChatGPT is mostly nontechnical people using the web UI, not developers using the API.

It's very possible the web UI is using a nerfed version of the model evident by its different versioning, but not the API which has more distinct versioning.

Post reply on HN