Earlier quoted context omitted.
I find that o1 and Sonnet 3.5 are good and bad quite equally on different things. That's why I keep asking both the same coding questions.
We do the same (all requests go to o1, sonnet and gemini and we store the results for later to compare) automatically for our research: Claude always wins. Even with specific prompting on both platforms. Especially frontend it seems o1 really is terrible.
GPT-5 is behind schedule
241–250 of 1001 posts
Re: GPT-5 is behind schedule
#242Earlier quoted context omitted.
No, I'm complaining that just because GPT-4 is called GPT-4 doesn't mean it's the fourth LLM from OpenAI. Off the top of my head: GPT-2, Codex, GPT-3 in three different flavors (babbage, curie, davinci), GPT-3.5. Suggesting that GPT-4 was "fourth" simply isn't credible. Just the other day they announced a jump from o1 to o3, skipping o2 purely because it's already the name of a major telecommunications brand in Europ…
It’s somehow funny to hear a British company being described as ‘in Europe’, but I suppose you’re technically correct…
Re: GPT-5 is behind schedule
#243Earlier quoted context omitted.
Is it "eerie"? LeCun has been talking about it for some time, and may also be OpenAI's rumored q-star, mentioned shortly after Noam Brown (diplomacybot) joining OpenAI. You can't hill climb tokens, but you can climb manifolds.
> You can't hill climb tokens, but you can climb manifolds. Could you explain this a bit please?
Title: "Objective Driven AI: Towards Machines that can Learn, Reason, and Plan"
Lytle Lecture Page: https://ece.uw.edu/news-events/lytle-lecture-series/
Slides: https://drive.google.com/file/d/1e6EtQPQMCreP3pwi5E9kKRsVs2N...
Re: GPT-5 is behind schedule
#244The lack of tech literacy in this article is a bit concerning: >Some researchers take this so seriously they won’t work on planes, coffee shops or anyplace where someone could peer over their shoulder and catch a glimpse of their work. I'm almost certain that originally this was meant to be a reference to public wifi networks, as planes and coffee shops are often the frequently cited prototypical examples. They made…
Re: GPT-5 is behind schedule
#245Earlier quoted context omitted.
It’s the same for me. I genuinely don’t understand how I can be having such a completely different experience from the people who rave about ChatGPT. Every time I’ve tried it’s been useless. How can some people think it’s amazing and has completely changed how they work, while for me it makes mistakes that a static analyser would catch? It’s not like I’m doing anything remarkable, for the past couple of months I’ve b…
> How can some people think it’s amazing and has completely changed how they work, while for me it makes mistakes that should a static analyser would catch? There are a lot of code monkeys working on boilerplate code, these people used to rely on stack overflow and now that chatgpt is here it's a huge improvement for them If you work on anything remotely complex or which hasn't been solved 10 times on stack overflow…
- write cvxpy code to find the chromatic number of a graph, and an optimal coloring, given its adjecency matrix.
- given an adjecency matrix write numpy code that enumerates all triangle-free vertex subsets.
- please port this old code from tensorflow to pytorch: ...
- in pytorch, i'd like to code a tensor network defining a 3-tensor of shape (d, d, d). my tensor consists of first projecting all three of its d-dimensional inputs to a k-dimensional vector, typically k=d/10, and then applying a (k, k, k) 3-tensor to contract these to a single number.
All were solved by ChatGPT on the first try.
Re: GPT-5 is behind schedule
#246I'm sure all the people who said "show me AI progress is slowing down!" 6 months ago will be acknowledging this article.
Re: GPT-5 is behind schedule
#247One fundamental challenge to me is that if each training run because more and more expensive, the time it takes it to learn what works/doesn't work widens. Half a billion dollars for training a model is already nuts, but if it takes 100 iterations to perfect it, you've cumulatively spent 50 billion dollars... Smaller models may actually be where rapid innovation continues simply because of tighter feedback loops. O3…
Re: GPT-5 is behind schedule
#248What we can reasonably assume from statements made by insiders: They want a 10x improvement from scaling and a 10x improvement from data and algorithmic changes The sources of public data are essentially tapped Algorithmic changes will be an unknown to us until they release, but from published research this remains a steady source of improvement Scaling seems to stall if data is limited So with all of that taken toge…
"With o3 now public knowledge, imagine how long it’s been churning out new thinking at expert level across every field." I highly doubt that. o3 is many orders of magnitude more expensive than paying subject matter experts to create new data. It just doesn't make sense to pay six figures in compute to get o3 to make data a human could make for a few hundred dollars.
> The process is painfully slow. GPT-4 was trained on an estimated 13 trillion tokens. A thousand people writing 5,000 words a day would take months to produce a billion tokens.
And if the human-generated data was so qualitatively good that it is smaller by three order of magnitudes, than I can assume it would be at least as expensive as o3.
Re: GPT-5 is behind schedule
#249Earlier quoted context omitted.
"With o3 now public knowledge, imagine how long it’s been churning out new thinking at expert level across every field." I highly doubt that. o3 is many orders of magnitude more expensive than paying subject matter experts to create new data. It just doesn't make sense to pay six figures in compute to get o3 to make data a human could make for a few hundred dollars.
Seems to me o3 prices would be what the consumer pays, not what OpenAI pays. That would mean o3 could be more efficient in-house than paying subject-matter experts.
In other words if you are diligent enough, you should at least validate your o3 solution with an actual expert for some time. You wouldn't just blindly trust OpenAI your business critical processes, would you? I would expect at least 3 month - 6 months for large corps and even more considering change management, re-upskilling, etc.
With all those considerations I really don't see the value prop at those prices and in those situations right now. Maybe if costs decrease ~1-3 orders of magnitude more for o3-low, depending on the the processes being automated.
Re: GPT-5 is behind schedule
#250Earlier quoted context omitted.
It’s somehow funny to hear a British company being described as ‘in Europe’, but I suppose you’re technically correct…
Technically…? Does anyone here believe that the EU and Europe is the same thing? Would you find it weird if someone said that a Norwegian company was in Europe?
I was just commenting on the fact that in the UK, ‘Europe’ generally means ‘continental Europe’.
> Would you find it weird if someone said that a Norwegian company was in Europe?
I’d find it weird if a European did. But from Americans it’s to be expected.