I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…
GPT-6 Astra
361–370 of 1001 posts
Re: GPT-6 Astra
#362The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
(/s, cause you never know these days)
[1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...
Re: GPT-6 Astra
#363Re: GPT-6 Astra
#364I'm going to call it. By 2030 all software is done and complete. But we are going to have more and new jobs.
'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
Re: GPT-6 Astra
#365Where is the cure for cancer?
Re: GPT-6 Astra
#366Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can i…
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
Re: GPT-6 Astra
#367Re: GPT-6 Astra
#368Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs. It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can i…
That's how cutting edge tech has always worked. Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.
Re: GPT-6 Astra
#369The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?
I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
Re: GPT-6 Astra
#370The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.