Live data from Hacker News

GPT-4.5

openai.com

631–640 of 1001 posts

Re: GPT-4.5

#631
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

For comparison, 3 years ago, the most powerful model out there (GPT-3 davinci) was $60/MTok.

Re: GPT-4.5

#632
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

It's crazy expensive because they want to pull in as much revenue as possible as fast as possible before the Open Source models put them outta business.

Re: GPT-4.5

#634

Earlier quoted context omitted.

In general yes, bench mark pollution is a big problem and why only dynamic benchmarks matter.

This is true, but how would pollution work for a benchmark designed to test hallucinations?

A dataset of labelled answers that are hallucinations and not hallucinations are published based on the benchmark as part of a paper.

People _seriously_ underestimate just how much stuff is online and how much impact it can have on training.

Re: GPT-4.5

#635
post #448

Earlier quoted context omitted.

It’d be great if someone would do that with the same data and prompt to other models. I did like the formatting and attributions but didn’t necessarily want attributions like that for every section. I’m also not sure if it’s fully matching what I’m seeing in the thread but maybe the data I’m seeing is just newer.

Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…

I noticed 4o mini didn't follow the directions to quote users. My favourite part of the 4.5 summary was how it quoted Antirez. 4o mini brought out the same quote, but failed to attribute it as instructed.

Re: GPT-4.5

#636

Earlier quoted context omitted.

I think that's the right interpretation, but that's pretty weak for a company that's nominally worth $150B but is currently bleeding money at a crazy clip. "We spent years and billions of dollars to come up with something that's 1) very expensive, and 2) possibly better under some circumstances than some of the alternatives." There are basically free, equally good competitors to all of their products, and pretty much…

I don’t mean to disagree too strongly, but just to illustrate another perspective: I don’t feel this is a weak result. Consider if you built a new version that you _thought_ would perform much better, and then you found that it offered marginal-but-not-amazing improvement over the previous version. It’s likely that you will keep iterating. But in the meantime what do you do with your marginal performance gain? Do you…

I've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.

Re: GPT-4.5

#637
I love the “listen to this article” widget doing embedded TTS for the article. Bugs / feedback:

The first words I hear are “introducing gee pee four five”. The TTS model starts cold? The next occurrence of the product name works properly as “gee pee tee four point five” but that first one in the title is mangled. Some kind of custom dictionary would help here too, for when your model needs to nail crucial phrases like your business name and your product.

No way of seeking back and forth (Safari, iOS 17.6.1). I don’t even need to seek, just replay the last 15s.

Very much need to select different voice models. Chirpy “All new Modern Family coming up 8/9c!” voice just doesn’t cut it for a science broadcast, and localizing models — even if it’s still English — would be even better. I need to hear this announcement in Bret Taylor voice, not Groupon CMO voice. (Sorry if this is your voice btw, and you work at OpenAI brandi. No offence intended.)

Re: GPT-4.5

#638

Earlier quoted context omitted.

You’re getting downvoted because you’re giving the same kind of hysterical reaction everyone derides crypto bros for. You also lead with the pretty strong assertion that previous commenter was lying, seemingly without proving proof anyone else can find.

It's directly from the post! I can't provide images here. I provided the numbers. What more can I do to show them? :)

People being wrong (especially on the internet) doesn't mean they are lying. Lying is being wrong intentionally.

Also, the person you replied to comments on the wording tricks they use. After suddenly bringing new data and direction in the discussion, even calling them "wrong" would have been a stretch.

I kindly suggest that you (and we all!) to keep discussing with an assumption of good faith.

Re: GPT-4.5

#639

Earlier quoted context omitted.

I don’t mean to disagree too strongly, but just to illustrate another perspective: I don’t feel this is a weak result. Consider if you built a new version that you _thought_ would perform much better, and then you found that it offered marginal-but-not-amazing improvement over the previous version. It’s likely that you will keep iterating. But in the meantime what do you do with your marginal performance gain? Do you…

I've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.

Maybe accidental, but I feel you’ve presented a straw man. We’re not discussing something that _may be_ better. It _is_ better. It’s not as big an improvement as previous iterations have been, but it’s still improvement. My claim is that reasonable people might still ship it.

Re: GPT-4.5

#640

Earlier quoted context omitted.

[flagged]

- Parent is still the top comment. - 2 hours in, -3. 2 replies: - [It's because] you're hysterical - [It's because you sound] like a crypto bro - [It's because] you make an equally unfounded claim - [It's because] you didn't provide any proof (Ed.: It is right in the link ! I gave the #s! I can't ctrl-F...What else can I do here...AFAIK can't link images...whatever, here's imgur. https://imgur.com/a/mkDxe78 ) - [It's…

Your original comment opened with:

  You are lying.
This is an ad hominem which assumes intent unknown to anyone other than the person to whom you replied.

Subsequently railing against comment rankings and enumerating curt summaries of other comments does not help either.

Post reply on HN