Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5 Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
GPT-6 Astra
301–310 of 1001 posts
Re: GPT-6 Astra
#302Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5 Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
Re: GPT-6 Astra
#303Re: GPT-6 Astra
#304The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new m…
Re: GPT-6 Astra
#305This is wild: OpenAI is basically declaring that AGI is here. https://www.theverge.com/ai-artificial-intelligence/989601/o... “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added,…
Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.
Re: GPT-6 Astra
#306I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…
Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
Re: GPT-6 Astra
#307Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
Re: GPT-6 Astra
#308I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…
I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).