Live data from Hacker News

GPT-4.5

openai.com

731–740 of 1001 posts

Re: GPT-4.5

#731
post #448

Earlier quoted context omitted.

It’d be great if someone would do that with the same data and prompt to other models. I did like the formatting and attributions but didn’t necessarily want attributions like that for every section. I’m also not sure if it’s fully matching what I’m seeing in the thread but maybe the data I’m seeing is just newer.

Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…

Compared to GPT-4.5 I prefer the GPT-4o version because it is less wordy. It summarizes and gives the gist of the conversation rather than reproducing it along with commentary.

Re: GPT-4.5

#732
post #635
post #448

Earlier quoted context omitted.

Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…

I noticed 4o mini didn't follow the directions to quote users. My favourite part of the 4.5 summary was how it quoted Antirez. 4o mini brought out the same quote, but failed to attribute it as instructed.

It's fascinating, but while this does mean it strays from the given example, I actually feel the result is a better summary. The 4.5 version is so long you might just read the whole thread yourself.

Re: GPT-4.5

#733
post #263

Earlier quoted context omitted.

Maybe if they build a few more data centers, they'll be able to construct their machine god. Just a few more dedicated power plants, a lake or two, a few hundred billion more and they'll crack this thing wide open. And maybe Tesla is going to deliver truly full self driving tech any day now. And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens. It's very difficult to r…

You have it all wrong. The end game is a scalable, reliable AI work force capable of finishing Star Citizen. At least this is the benchmark for super-human general intelligence that I propose.

I’m surprised ‘create superhuman agi’ isn’t a stretch goal on their everlasting funding drive. Seems like a perfect Robertsian detour.

Re: GPT-4.5

#734

In many ways I'm not an OpenAI fan (but I need to recognize their many merits). At the same time, I believe people are missing what they tried to do with GPT 4.5: it was needed and important to explore the pre-training scaling law in that direction. A gift to science, however selfist it could be.

> A gift to science This is hardly recognizable as science. edit: Sorry, didn't feel this was a controversial opinion. What I meant to say was that for so-called science, this is not reproducible in any way whatsoever. Further, this page in particular has all the hallmarks of _marketing_ copy, not science. Sometimes a failure is just a failure, not necessarily a gift. People could tell scaling wasn't working well bef…

> People could tell scaling wasn't working well before the release of GPT 4.5.

Different people tell different things all the time. That's not science. Experiment is science.

Re: GPT-4.5

#735
post #426
post #384

Earlier quoted context omitted.

most humans are generally intelligent but can't do what you just said AGI should do...

Excluding the realtime-iness, humans do at least possess the capacity to do so. Besides, humans are capable of rigorous logic (which I believe is the most crucial aspect of intelligence) which I don’t think an agent without a proof system can do.

yes the problem is that there is no consensus about what AGI should be: https://medium.com/@fsndzomga/there-will-be-no-agi-d9be9af44...

Re: GPT-4.5

#736
post #66

Finally a scaling wall? This is apparently (based on pricing) using about an order of magnitude more compute, and is only maybe 10% more intelligent. Ideally DeepSeeks optimizations help bring the costs way down, but do any AI researchers want to comment on if this changes the overall shape of the scaling curve?

It depends on how you compare.

On a subset of tasks I'm interested in, it's 10x more intelligent than GPT-4. (Note that GPT-4 was in many ways better than 4o.)

It's not a coding champion, but it knows A LOT of stuff, excellent common sense, top quality writing. For me it's like "deep research lite".

I found OpenAI Deep research excellent, but GPT-4.5 might in many cases beat it.

Re: GPT-4.5

#737

"Starting today, ChatGPT Pro users will be able to select GPT‑4.5 in the model picker on web, mobile, and desktop. We will begin rolling out to Plus and Team users next week, then to Enterprise and Edu users the following week." Thanks for being transparent about this. Nothing is more frustrating than being locked out for indeterminate time from the hot thing everyone talks about. I hope the announcement is true with…

I'm outside the US and I have access to ChatGPT 4.5 with ChatGPT Pro subscription. Didn't have that access yesterday at the time of announce, but probably they were staggering the release a bit to even the load over multiple hours.

Re: GPT-4.5

#738

Seeing OpenAI and Anthropic go different routes here is interesting. It is worth moving past the initial knee jerk reaction of this model being unimpressive and some of the comments about "they spent a massive amount of money and had to ship something for it..." * Anthropic appears to be making a bet that a single paradigm (reasoning) can create a model which is excellent for all use cases. * OpenAI seems to be betti…

Ever since those kids demo'd their fact checking engine here, which was just Input -> LLM -> Fact Database -> LLM -> LLM -> Output I have been betting that it will be advantageous to move in this general direction.

Re: GPT-4.5

#739
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…

> https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySddv...

how do you do the different size circles and colored sequences like that? this is god tier skills

Re: GPT-4.5

#740

Earlier quoted context omitted.

It still not smart enough to replace for example customer service.

It's absolutely able to replace the majority of customer service volume which is full of mundane questions.

Such brutal reductionism: how do you calculate an ever growing percentage of customers so pissed at this terrible service that you lose customers forever? Not just one company losing customers... but an entire population completely distrusting and pulling back from any and all companies pulling this trash
Post reply on HN