Live data from Hacker News

GPT-6 Astra

openai.com

611–620 of 1001 posts

Re: GPT-6 Astra

#611

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI). If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.

Is that hyperbole or do you know of a specific "AI website" that uses 100% of your CPU?

Re: GPT-6 Astra

#612

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI. It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big buc…

If they adopt a different tone (like Anthropic has been doing), it will get called fear marketing.

Re: GPT-6 Astra

#613
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!

Re: GPT-6 Astra

#614
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.

Same for HealthBench Professional and a few others.

Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.

Re: GPT-6 Astra

#615
post #402

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

I am going to take my 401k and open up a coffee or bike shop. If I am going to be broke, might as well enjoy what I do.

who's buying your coffee or bikes when the rest of people are broke? Neither of those is a survival necessity.

Re: GPT-6 Astra

#616
post #523

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks: https://artificialanalysis.ai/models Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.

And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.

Re: GPT-6 Astra

#617
post #56

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

What's just as interesting is this morning Axiom Math announced 212 and OpenAI then appears to have rushed out their 186 announcement just 1-2 hours later followed by Astra. Did they accelerate the release of Astra itself? Not necessarily, but it definitely looks like they ended up pushing much harder and faster than planned on their 186 result. X activity suggests Anthropic had a similar result as well but wasn't as fast as OpenAI in packaging it up and sharing it in response to Axiom, so they mostly just bolted onto OpenAI's messaging.

The reason I think this is interesting is that Axiom is a tiny lab in comparison that wouldn't have had access to Astra at all. I'd be curious to learn how Axiom is able to effectively compete at this frontier with vastly fewer resources.

Re: GPT-6 Astra

#618

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.

Re: GPT-6 Astra

#619

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

I have exactly the same thoughts - or perhaps slightly bleaker ones - evry time I read this relentless stream of news about new model releases. I’m tired of all the enthusiastic comments about how excited everyone is about the latest benchmark results and so on. I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just si…

Are you doing less work now because of AI? Not sure about you, but I'm doing a lot more work. Not saying I like doing more, or how it's getting done, but nonetheless I don't feel like AI is doing what the CEOs of AI labs want to convince everyone of.

Re: GPT-6 Astra

#620

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?

Because some people do.

Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.

Post reply on HN