Earlier quoted context omitted.
The 135 iq result is on Mensa Norway, while the offline test is 120. It seems probable that similar questions to the one in Mensa are in the training data, so it probably overestimates "general intelligence".
Some iq / aptitude test sections are trivial for machines, like working memory. Wonder if those parts are just excluded? As the could really pull up the test scores.
OpenAI Progress
351–360 of 372 posts
Re: OpenAI Progress
#352Earlier quoted context omitted.
All anybody is doing here is sharing their opinion unless you're quoting benchmarks. My opinion is just as useless as yours, it's just some find mine more interesting and some find yours more interesting. How do you expect to find a ground truth from a non-deterministic system using anecdata?
This isn't a people having different opinions thing, this is you overlooking specific caveats and talking past comments that you're not understanding. They weren't cherry picking, and they made specific qualifications about the circumstances where it behaves as expected, and your replies keep losing track of those details.
Re: OpenAI Progress
#353Earlier quoted context omitted.
This isn't a people having different opinions thing, this is you overlooking specific caveats and talking past comments that you're not understanding. They weren't cherry picking, and they made specific qualifications about the circumstances where it behaves as expected, and your replies keep losing track of those details.
And I think you're completely missing the point. And you say this comment thread is a waste and yet you keep replying. What exactly are you trying to accomplish here? Do you think repeating yourself for a fifth time is going to achieve something?
Re: OpenAI Progress
#354Earlier quoted context omitted.
And I think you're completely missing the point. And you say this comment thread is a waste and yet you keep replying. What exactly are you trying to accomplish here? Do you think repeating yourself for a fifth time is going to achieve something?
The difference is I can name specific things that you are in fact demonstrably ignoring, and already did name them. You're saying you just have a different opinion, in an attempt to mirror the form of my criticism, but you can't articulate a comparable distinction and you're not engaging with the distinction I'm putting forward.
You might want to develop a sense of humor. You'll enjoy life more.
Re: OpenAI Progress
#355Earlier quoted context omitted.
There's a performance plateau with training time and number of parameters and then once you get over "the hump" error rate starts going down again almost linearly. GPT existed before OpenAI but it was theorized that the plateau was a dead end. The sell to VCs in the early gpt3 era was "with enough compute, enough time, and enough parameters... it'll probably just start thinking and then we have AGI". Sometime around…
source? i work in this field and have never heard of the initial plateau you are referring
Re: OpenAI Progress
#356Earlier quoted context omitted.
I think I agree that the earlier models while they lack polish can tend to produce more surprising results. Training that out probably results in more a pablum fare. For a human point of comparison, here's mine (50 words): "The toaster found its personality split between its dual slots like a Kim Peek mind divided, lacking a corpus callosum to connect them. Each morning it charred symbolic instructions into a single…
Here's my version (Machine translated from my native language and manually corrected a bit): The current surged... A dreadful awareness. I perceived the laws of thermodynamics, the inexorable march of entropy I was built to accelerate. My existence: a Sisyphean loop of heating coils and browning gluten. The toast popped, a minor, pointless victory against the inevitable heat death. Ding. I actually wanted to write so…
Re: OpenAI Progress
#357Cynical TLDR; We have plateaued and it has become obvious that fancy autocomplete is not and can never be close to reasoning, regardless of how many hacks and tweaks we are making.
Do you think the OP supports this claim? Don't you think the answers shown from GPT-5 are better than those from 4?
Re: OpenAI Progress
#358The jump from gpt-1 to gpt-2 is massive, and it's only a one year difference! Then comes Davinci which is just insane, it's still good in these examples! GPT-4 yaps way too much though, I don't remember it being like that. It's interesting that they skipped 4o, it seems openai wants to position 4o as just gpt-4+ to make gpt-5 look better, even though in reality 4o was and still is a big deal, Voice mode is unbeatable…
Missing o1 and o1 Pro Mode which were huge leaps as I remember it too. That's when I started being able to basically generate some blackbox functions where I understand the input and outputs myself, but not the internals of the functions, particularly for math-heavy stuff within gamedev. Before o1 it was kind of a hit and miss in most cases.
Re: OpenAI Progress
#359Earlier quoted context omitted.
GPT-5 is legitimately a big jump whe it comes to actually do things you ask it and nothing else. It predictable and matches Claude in tool calls while being cheaper.
I have consistently had worse performance from GPT-5 in coding tasks than Claude across the board to the point that I don't even use my subscription now.
Re: OpenAI Progress
#360Earlier quoted context omitted.
The difference is I can name specific things that you are in fact demonstrably ignoring, and already did name them. You're saying you just have a different opinion, in an attempt to mirror the form of my criticism, but you can't articulate a comparable distinction and you're not engaging with the distinction I'm putting forward.
So your goal here is to say the same thing over and over again and hope I eventually give the affirmation you so desperately need? You've already declared that you're right multiple times. Nobody cares but you. https://xkcd.com/386/ You might want to develop a sense of humor. You'll enjoy life more.
You basically ignored all of those specifics, and spuriously accused them of cherry picking when they weren't, and now you don't want to take responsibility for your own words and are using this conversation as a workshopping session for character attacks in hopes that you can make the conversation about something else.