The lack of tech literacy in this article is a bit concerning: >Some researchers take this so seriously they won’t work on planes, coffee shops or anyplace where someone could peer over their shoulder and catch a glimpse of their work. I'm almost certain that originally this was meant to be a reference to public wifi networks, as planes and coffee shops are often the frequently cited prototypical examples. They made…
GPT-5 is behind schedule
221–230 of 1001 posts
Re: GPT-5 is behind schedule
#222Earlier quoted context omitted.
I think the wildest thing is actually Meta’s latest paper where they show a method for LLMs reasoning not in English, but in latent space https://arxiv.org/pdf/2412.06769 I’ve done research myself adjacent to this (mapping parts of a latent space onto a manifold), but this is a bit eerie, even to me.
Is it "eerie"? LeCun has been talking about it for some time, and may also be OpenAI's rumored q-star, mentioned shortly after Noam Brown (diplomacybot) joining OpenAI. You can't hill climb tokens, but you can climb manifolds.
Could you explain this a bit please?
Re: GPT-5 is behind schedule
#223Earlier quoted context omitted.
No, I'm complaining that just because GPT-4 is called GPT-4 doesn't mean it's the fourth LLM from OpenAI. Off the top of my head: GPT-2, Codex, GPT-3 in three different flavors (babbage, curie, davinci), GPT-3.5. Suggesting that GPT-4 was "fourth" simply isn't credible. Just the other day they announced a jump from o1 to o3, skipping o2 purely because it's already the name of a major telecommunications brand in Europ…
While I’m sure it’s unintentional, that amounts to nitpicking. I can easily find three to include and pass over the rest. Face value turns out to be a decent approximation.
Re: GPT-5 is behind schedule
#224Earlier quoted context omitted.
If you are talking about ARC benchmark, then o3-low doesn't look that special if you take into account there are plenty of finetuned models with much smaller resources achieved 40-50% results on private set (not semi-private like o3-low).
- I'm not just talking about ARC. On frontier Math, we have 2 scores, one with pass@1 and another with consensus vote with 64 samples. Both scores are much better than previous Sota. - Also apparently, ARC wasn't a special fine-tune but rather some of the training set in the corpus for pre-training.
that result is not verifiable, not reproducable, unknown if it was leaked and how it was measured. Its kinda hype science.
> ARC wasn't a special fine-tune but rather some of the training set in the corpus for pre-training.
post says: Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details.
So, I guess we don't know.
Re: GPT-5 is behind schedule
#225Earlier quoted context omitted.
I think my biggest pet peeve is when someone shares an insight which is unmistakably based on intuition, inference, critical thinking, etc (all mental faculties we are allowed to use to come to conclusions in the face of information asymmetry btw) ...and then gets hit deadpan with the good old "Source?", like it's some sort of gotcha. I think people have started to confuse "making logical conclusions without perfect…
I think it's okay to make logical conclusions but you must base them in evidence, not just suppositions. Intuition is a good start to begin generating hypothesis, but it doesn't render conclusions. I interpreted the GP asking for sources as "can you give me some evidence that would help me reach the same conclusions you've reached". I think that's much preferable to just accepting random things people say at face val…
My point is simply that is we can skip the passive aggressiveness and just say "can you give me some more evidence that would help me reach the same conclusions you've reached".
Otherwise you're not actually asking for a source, you're just saying "I disagree" in a very roundabout way.
Re: GPT-5 is behind schedule
#226Earlier quoted context omitted.
It's reasonable to ask for sources when an opinion is phrased as a fact, as GGP did. I don't see how you got that it was _unmistakably_ an opinion from that comment. There is no way to deduce by intuition alone that GPT-5 == GPT-4o. So either that person has some information the rest of us aren't privy to, or it's an opinion phrased as a fact. In either case, it deserves clarification.
On a second read I see that the comment notes that it is intended as speculation, but still it seems rather confident in its own accuracy and I am not even sure it's wrong, but just looking for something that warrants the confidence.
Looking at the benchmarks it was also very expected in my opinion. Sure, the absolute results are/were sky high, but results relative to the previous gen were not exponential now, they were comparatively smaller than between 2 and 3, or 3 and 4. So I'm guessing that they have invested and worked for 2023-2024 on a brand new model, and branded it according to the model results.
Re: GPT-5 is behind schedule
#227Earlier quoted context omitted.
It doesn't even look like 4o is scaled up parameter wise from 4 and was released closer in time than either 3 or 4 were from their predecessors at a time where the scaling required for these next gen iterations has only gotten more difficult. Critical thinking ? Lol it's just blind speculation.
If you disagree with their reasoning then you explain that . You don't do this passive aggressive "source???" thing. It's a bit like starting a Slack conversation with "Hi?": we all know you have a secondary objective, but now you're inserting an extra turn of phrase into the mix
To me, OP's speculation reads as obvious nonsense but that might not be the case for everybody. Asking for sources or such to what is entirely speculation is perfectly valid and personally, that comment does not ring as passive aggressive to me but maybe it's just me.
Just because someone doesn't know enough to refute the reasoning doesn't mean they must take whatever they read at face value.
Re: GPT-5 is behind schedule
#228Earlier quoted context omitted.
While I’m sure it’s unintentional, that amounts to nitpicking. I can easily find three to include and pass over the rest. Face value turns out to be a decent approximation.
If this was a random blog post I wouldn't nitpick, but this is the Wall Street Journal.
Re: GPT-5 is behind schedule
#229Re: GPT-5 is behind schedule
#230Earlier quoted context omitted.
If you disagree with their reasoning then you explain that . You don't do this passive aggressive "source???" thing. It's a bit like starting a Slack conversation with "Hi?": we all know you have a secondary objective, but now you're inserting an extra turn of phrase into the mix
Not everyone keeps up with LLM development enough to know how far apart the release dates for these models are, how much scaling (roughly) has been done on each iteration and a decent ballpark for how much open ai might try to scale up a next gen model. To me, OP's speculation reads as obvious nonsense but that might not be the case for everybody. Asking for sources or such to what is entirely speculation is perfectl…
If anything just breezily asking for a source would imply to people who don't know better that this is a rather even keeled take and just needs some more evidence on top. "I disagree and here's why" nips that in the bud directly.