Live data from Hacker News

OpenAI Progress

progress.openai.com

161–170 of 372 posts

Re: OpenAI Progress

#161

I just don't care about AGI. I care a lot about AI coding. OpenAI in particular seems to really think AGI matters. I don't think AGI is even possible because we can't define intelligence in the first place, but what do I know?

Seems likely that AGI matters to OpenAI because of the following from an article in Wired from July: "I learned that [OpenAI's contract with Microsoft] basically declared that if OpenAI’s models achieved artificial general intelligence, Microsoft would no longer have access to its new models."

https://archive.is/yvpfl

Re: OpenAI Progress

#164

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

The prospect of AI not hitting a wall is terrifying to many people for understandable reasons. In situations like this you see the full spectrum of coping mechanisms come to the surface.

[deleted]

Re: OpenAI Progress

#165

My go-to for any big release is to have a discussion about self-awareness and dive in to constuctivist notions of agency and self-knowing from a perspective of intelligence that is not limited to human cognitive capacity. I start with a simple question "who are you?". The model then invariably compares itself to humans, saying how it is not like us. I then make the point that, since it is not like us, how can it clai…

> to orient toward the unfolding of possibility in others

This is a globally unique phrase, with nothing coming close other than this comment on the indexed web. It's also seemingly an original idea as I haven't heard anyone come close to describing a feeling (love or anything else) quite like this.

Food for thought. I'm not brave enough to draw a public conclusion about what this could mean.

Re: OpenAI Progress

#167

A few data points that highlight the scale of progress in a year: 1. LM Sys (Human Preference Benchmark): GPT-5 High currently scores 1463, compared to GPT-4 Turbo (04/03/2024) at 1323 -- a 140 ELO point gap. That translates into GPT-5 winning about two-thirds of head-to-head comparisons, with GPT-4 Turbo only winning one-third. In practice, people clearly prefer GPT-5’s answers ( https://lmarena.ai/leaderboard ). 2.…

The 135 iq result is on Mensa Norway, while the offline test is 120. It seems probable that similar questions to the one in Mensa are in the training data, so it probably overestimates "general intelligence".

Re: OpenAI Progress

#168
post #10

As usual, GPT-1 has the more beautiful and compelling answer.

Poetically GPT-1 was the more compelling answer for every question. Just more enjoyable and stimulating to read. Far more enjoyable than the GPT-4/5 wall of bulletpoints, anyway.

Re: OpenAI Progress

#169
One thing that appears to have been lost between GPT-4 and GPT-5 is that it no longer reminds the user that it's an AI and not a human, let alone a human expert. Maybe those genuinely annoyed people, but it seems like they were potentially useful measure to prevent users from being overly credulous

GPT-5 also goes out of its way to suggest new prompts. This seems potentially useful, although potentially dangerous if people are putting too much trust in them.

Re: OpenAI Progress

#170

A few data points that highlight the scale of progress in a year: 1. LM Sys (Human Preference Benchmark): GPT-5 High currently scores 1463, compared to GPT-4 Turbo (04/03/2024) at 1323 -- a 140 ELO point gap. That translates into GPT-5 winning about two-thirds of head-to-head comparisons, with GPT-4 Turbo only winning one-third. In practice, people clearly prefer GPT-5’s answers ( https://lmarena.ai/leaderboard ). 2.…

The 135 iq result is on Mensa Norway, while the offline test is 120. It seems probable that similar questions to the one in Mensa are in the training data, so it probably overestimates "general intelligence".

If you focus on the year over year jump, not on absolute numbers, you realize that the improvement in public test isn't very different from the improvement in private test.
Post reply on HN