Live data from Hacker News

GPT-5.3-Codex

openai.com

441–450 of 634 posts

Re: GPT-5.3-Codex

#441
post #352

Earlier quoted context omitted.

Moltbot is an attempt to do that. Would you hire it as a personal secretary and entrust all your personal data to it?

Only people who haven't had a secretary would think it's a personal secretary. Like, it can't even answer the phone.

There are plenty of companies that sell an AI assistant that answers the phone as a service, they just aren't named OpenAI or Anthropic. They'll let callers book an appointment onto your calendar, even!

Re: GPT-5.3-Codex

#442

Earlier quoted context omitted.

Is an imperfect yardstick better than no yardstick? It reminds me of documentation — the only thing worse than no documentation is wrong documentation.

Yes, because there’s value in a common reference for comparison. It helps to shed light on different models’ relative strengths and weaknesses. And, just like with performance benchmarks, you can learn to spot and read past the ways that people game their results. The danger is really more in when people who are less versed in the subject matter take what are ultimately just a semi tamed genre of sales pitch at face…

Thanks, that makes sense!

Re: GPT-5.3-Codex

#443
post #178

Earlier quoted context omitted.

I just wanted to say that's a pretty cool demo! I hadn't realised people were using it for things like this.

Thank you. There's a demo save to get the full feel of it quickly. There's also a 2D-ASCII and 3D render you can hotswap between. The 3D models are generated with Meshy. The entire game is 'AI slop'. I intentionally did no code reviews to see where that would get me. Some prompts were very specific but other prompts were just 'add a research of your choice'. This was built using old versions of Codex, Gemini and Clau…

Any estiimates on how much it cost you? In terms of total real world time, money, and time spent by the agents.

Re: GPT-5.3-Codex

#444

Earlier quoted context omitted.

Yes, but also you'll never have any early evidence of the Foom until the Foom itself happens.

If only General Relativity had such an ironclad defense of being as unfalsifiable as Foom Hypothesis is. We could’ve avoided all of the quantum physics nonsense.

it doesn't mean it's unfalsifiable - it's a prediction about the future so you can falsify it when there's a bound on when it is going to happen. it just means there's little to no warning. I think it's a significant risk to AI progress that it can reach some sort of improvement speed > speed of warning or any threats from AI improvement

Re: GPT-5.3-Codex

#445

Earlier quoted context omitted.

More importantly, this is the early steps of a model self improving itself. Do we still think we'll have soft take off?

> Do we still think we'll have soft take off? There's still no evidence we'll have any take off. At least in the "Foom!" sense of LLMs independently improving themselves iteratively to substantial new levels being reliably sustained over many generations. To be clear, I think LLMs are valuable and will continue to significantly improve. But self-sustaining runaway positive feedback loops delivering exponential improv…

To me FOOM means like the hardest of hard takeoffs and improving at a sustained rate which is higher than without humans is not a takeoff at all.

Re: GPT-5.3-Codex

#446

Earlier quoted context omitted.

This has already been going on for years. It's just that they were using GPT 4.5 to work on GPT 5. All this announcement mean is that they're confident enough in early GPT 5.3 model output to further refine GPT 5.3 based on initial 5.3. But yes, takeoff will still happen because of this recursive self improvement works, it's just that we're already past the inception point.

I can't tell if this is a serious conversation anymore.

I think it's important in AI discussions to reason correctly from fundamentals and not disregard possibilities simply because they seem like fiction/absurd. If the reasoning is sound, it could well happen.

Re: GPT-5.3-Codex

#447

Earlier quoted context omitted.

Sounds like the researchers behind https://ai-2027.com/ haven't been too far off so far.

We'll see. The first two things that they said would move from "emerging tech" to "currently exists" by April 2026 are: - "Someone you know has an AI boyfriend" - "Generalist agent AIs that can function as a personal secretary" I'd be curious how many people know someone that is sincerely in a relationship with an AI. And also I'd love to know anyone that has honestly replaced their human assistant / secretary with a…

It's important to remember though (this is besides the point for what you're saying) that job displacement of things like secretaries from AI do not require it to be a near perfect replacement. There are many other factors (for example if it's much cheaper and can do part of the work it can dramatically shrink demand as people can shift to an imperfect replacement in AI)

Re: GPT-5.3-Codex

#448
post #206

Earlier quoted context omitted.

> With Codex (5.3), the framing is an interactive collaborator: you steer it mid-execution, stay in the loop, course-correct as it works. > With Opus 4.6, the emphasis is the opposite: a more autonomous, agentic, thoughtful system that plans deeply, runs longer, and asks less of the human. Ain't the UX is the exact opposite? Codex thinks much longer before gives you back the answer.

I've also had the exact opposite experience with tone. Claude Code wants to build with me, and Codex wants to go off on its own for a while before returning with opinions.

well with the recent delays i can easily find claude code going off on it's own for 20 minutes and have no idea what it's going to come back with. but one time it overflowed it's context on a simple question, and then used up the rest of my session window. in a way a lot of ai assistants have ime have this awkward thing where they complicate something in a non-visible and think about it for a long time burning up context before coming up with a summary based upon some misconception.

Re: GPT-5.3-Codex

#450

Earlier quoted context omitted.

The consumers are getting huge wins. Model costs continue to collapse while capability improves. Competition is fantastic.

> The consumers are getting huge wins. However, the investors currently subsidizing those wins to below cost may be getting huge losses.

Yes, but that's the nature of the game, and they know it.
Post reply on HN