Live data from Hacker News

GPT-6 Astra

openai.com

811–820 of 1001 posts

Re: GPT-6 Astra

#811

What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of a…

> How will we make a living? Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output. Think of the artists: I'm sorry for those too, for for many it's already game over today . > How will…

> Think of the artists: I'm sorry for those too, for for many it's already game over today.

No it's not??? People want human made art. Abstract artists didn't paint anything technically challenging and they sold art for millions because art is about human connection and human inventiveness, not whatever prompt you could feed into an AI.

> he'll also help your company fix the mess LLMs created.

Hugely valuable. I keep a list of every PR I halted which was AI generated with AI commit summaries and AI tests and no human reviewer till me. It's very easy to convince people to keep me around. AIs are smarter than ever but the people using them and putting the prompts in and not reading the output are still as dumb as before

Re: GPT-6 Astra

#812

That hero video is interesting. A projector and speech. Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now. The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for m…

given the fact that we've moved in my office from 3-people offices to open-plan office to flex desk now I'm not exactly sure I would want my coworkers to speak all day to their computers and gesturing / walking in front of a projector (provided there will still be coworkers with IA)

Re: GPT-6 Astra

#814

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

The rocket they 3D print at the end isn't even the same as the one in the game. It has a curved body, but the game asset is a cylinder.

Re: GPT-6 Astra

#815

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains.

Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wouldn’t necessarily say they are good at brand new situations, but there has definitely been progress.

However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level, let alone drive a robot or other non-language tasks. Its architecture and ability to learn seem a long way off from being general.

Vision models, being able to encompass language and much more, seem to me like a theoretically closer step to AGI. Yet, there is a lot more to the world than just what we can see.

On the other hand, in humans, vision certainly is not necessary for intelligence. So there is something more fundamental, neither vision nor language, that high levels of intelligence are based upon. Once we figure that out, I think we will be able to build AGI.

Re: GPT-6 Astra

#816

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate. "Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip! I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming…

And nobody with actual taste is going to take any of those recs seriously in the first place.

Re: GPT-6 Astra

#817

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains. Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wo…

> In my view, intelligence includes an ability to learn and adapt to never-before-seen situations.

This is exactly what ARC AGI tests

> And then general intelligence is an ability to apply that across a wide variety of domains.

My experience with Fable is that it can certainly apply that in a wide variety of domains

> However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level

People also struggle to play Chess at a basic level. They only succeed by studying the game for a long time. I will concede that humans can do this and LLMs generally cannot.

Re: GPT-6 Astra

#819

I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously? Even if I did trust an AI to get everything right, it's not like the AI can read my mind. If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really…

Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines. I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries

They’re ridiculously facile use cases for what agentic AI can already do.

I don’t trust an agent with full executive power - yet - but I am essentially using GPT as my PA. Right now I’ve got it managing a construction project with recalcitrant contractors, an international move and visa tied to a property purchase, a short term holiday rental business, and basically just popping up in my life going “the situation is this, you need to do X/I suggest you send Y to Z, the email is prepped in your drafts”. I am of course also using it for software development, and have had it resolve every digital chore in my home life.

I guess it doesn’t make for a quick elevator pitch, but I’m finding it has reduced my cognitive overhead on a whole raft of fuckery that would otherwise have me shouting at inanimate objects.

Anyway, today I have to take the cat to the vet, email a bank, go and discuss ceiling systems, and review a proposed shopping list for a maintenance visit to a holiday rental. Can’t wait for the bleeding robots to get here.

Re: GPT-6 Astra

#820

I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right. The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. It's a tough balance to get right, and although this has been possible to achieve with additional pro…

>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in th…

> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed.

That's exactly what i want to happen. I hate when it assumes my direct question was an indirect instruction

Post reply on HN