Earlier quoted context omitted.
It’d be great if someone would do that with the same data and prompt to other models. I did like the formatting and attributions but didn’t necessarily want attributions like that for every section. I’m also not sure if it’s fully matching what I’m seeing in the thread but maybe the data I’m seeing is just newer.
Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…
GPT-4.5
731–740 of 1001 posts
Re: GPT-4.5
#732Earlier quoted context omitted.
Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…
I noticed 4o mini didn't follow the directions to quote users. My favourite part of the 4.5 summary was how it quoted Antirez. 4o mini brought out the same quote, but failed to attribute it as instructed.
Re: GPT-4.5
#733Earlier quoted context omitted.
Maybe if they build a few more data centers, they'll be able to construct their machine god. Just a few more dedicated power plants, a lake or two, a few hundred billion more and they'll crack this thing wide open. And maybe Tesla is going to deliver truly full self driving tech any day now. And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens. It's very difficult to r…
You have it all wrong. The end game is a scalable, reliable AI work force capable of finishing Star Citizen. At least this is the benchmark for super-human general intelligence that I propose.
Re: GPT-4.5
#734In many ways I'm not an OpenAI fan (but I need to recognize their many merits). At the same time, I believe people are missing what they tried to do with GPT 4.5: it was needed and important to explore the pre-training scaling law in that direction. A gift to science, however selfist it could be.
> A gift to science This is hardly recognizable as science. edit: Sorry, didn't feel this was a controversial opinion. What I meant to say was that for so-called science, this is not reproducible in any way whatsoever. Further, this page in particular has all the hallmarks of _marketing_ copy, not science. Sometimes a failure is just a failure, not necessarily a gift. People could tell scaling wasn't working well bef…
Different people tell different things all the time. That's not science. Experiment is science.
Re: GPT-4.5
#735Earlier quoted context omitted.
most humans are generally intelligent but can't do what you just said AGI should do...
Excluding the realtime-iness, humans do at least possess the capacity to do so. Besides, humans are capable of rigorous logic (which I believe is the most crucial aspect of intelligence) which I don’t think an agent without a proof system can do.
Re: GPT-4.5
#736Finally a scaling wall? This is apparently (based on pricing) using about an order of magnitude more compute, and is only maybe 10% more intelligent. Ideally DeepSeeks optimizations help bring the costs way down, but do any AI researchers want to comment on if this changes the overall shape of the scaling curve?
On a subset of tasks I'm interested in, it's 10x more intelligent than GPT-4. (Note that GPT-4 was in many ways better than 4o.)
It's not a coding champion, but it knows A LOT of stuff, excellent common sense, top quality writing. For me it's like "deep research lite".
I found OpenAI Deep research excellent, but GPT-4.5 might in many cases beat it.
Re: GPT-4.5
#737"Starting today, ChatGPT Pro users will be able to select GPT‑4.5 in the model picker on web, mobile, and desktop. We will begin rolling out to Plus and Team users next week, then to Enterprise and Edu users the following week." Thanks for being transparent about this. Nothing is more frustrating than being locked out for indeterminate time from the hot thing everyone talks about. I hope the announcement is true with…
Re: GPT-4.5
#738Seeing OpenAI and Anthropic go different routes here is interesting. It is worth moving past the initial knee jerk reaction of this model being unimpressive and some of the comments about "they spent a massive amount of money and had to ship something for it..." * Anthropic appears to be making a bet that a single paradigm (reasoning) can create a model which is excellent for all use cases. * OpenAI seems to be betti…
Re: GPT-4.5
#739GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…
The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…
how do you do the different size circles and colored sequences like that? this is god tier skills
Re: GPT-4.5
#740Earlier quoted context omitted.
It still not smart enough to replace for example customer service.
It's absolutely able to replace the majority of customer service volume which is full of mundane questions.