Live data from Hacker News

ChatGPT agent: bridging research and action

openai.com

381–390 of 508 posts

Re: ChatGPT agent: bridging research and action

#381
post #311

Earlier quoted context omitted.

This is a big moving of the goalposts. The optimists were saying Level 5 would be purchasable everywhere by ~2018. They aren’t purchasable today, just hail-able. And there’s a lot of remote human intervention. And San Francisco doesn’t get snow.

Hell - SF doesn’t have motorcyclists or any vehicular traffic, driving on the wrong side of the road. Or cows sharing the thoroughfares. It should be obvious to all HNers that have lived or travelled to developing / global south regions - driving data is cultural data. You may as well say that self driving will only happen in countries where the local norms and driving culture is suitable to the task. A desperately a…

I'm in the Philippines now, and that's how I know this is the correct take. Especially this part:

"Driving data is cultural data."

The optimists underestimate a lot of things about self-driving cars.

The biggest one may be that in developing and global south regions, civil engineering, design, and planning are far, far away from being up to snuff to a level where Level 5 is even a slim possibility. Here on the island I'm on, the roads, storm water drainage (if it exists at all) and quality of the built environment in general is very poor.

Also, a lot of otherwise smart people think that the increment between Level 4 and Level 5 is the same as that between all six levels, when the jump from Level 4 to Level 5 automation is the biggest one and the hardest to successfully accomplish.

Re: ChatGPT agent: bridging research and action

#382
post #65

Very slightly impressed by their emphasis on the gigantic (my word, not theirs) risk of giving the thing access to real creds and sensitive info.

The sane way to do this (if you wanted to) would be to give the AI a debit card with a small balance to work with. If funds get stolen, you know exactly what the maximum damage is. And if you can't afford that damage, then you wouldn't have been able to afford that card to begin with. But since people can cancel transactions with a credit card, that's what people are going to do, and it will be a huge mess every time…

It's not like a credit card is all that different from a debit card in terms of cancellations. If this becomes a big enough problem, I would imagine that card issuers will simply stop accepting "my agent did it" as an excuse in chargeback requests.

Re: ChatGPT agent: bridging research and action

#384

Earlier quoted context omitted.

In the context of our conversation and what OP wrote, there has been no breakthrough since around 2018. What you're seeing is the harvesting of all low-hanging fruit from a tree that was discovered years ago. But fruit is almost gone. All top models perform at almost the same level. All the "agents" and "reasoning models" are just products of training data. I wrote more about it here: https://news.ycombinator.com/ite…

This "all breakthroughs are old" argument is very unsatisfying. It reminds me of when people would describe LLMs as being "just big math functions". It is technically correct, but it misses the point. AI researchers spent years figuring out how to apply RL to LLMs without degrading their general capabilities. That's the breakthrough. Not the existence of RL, but making it work for LLMs specifically. Saying "it's just…

I understand your argument. The recipe that finally let RLHF + SFT work without strip mining base knowledge was real R&D, and GPT 4 class models wouldn’t feel so "chatty but competent" without it. I just still see ceiling effects that make the whole effort look more like climbing a very tall tree than building a Saturn V.

GPT 4.1 is marketed as a "major improvement" but under the hood it’s still the KL-regularised PPO loop OpenAI first stabilized in 2022 only with a longer context window and a lot more GPUs for reward model inference.

They retired GPT 4.5 after five months and told developers to fall back to 4.1. The public story is "cost to serve” not breakthroughs left on the table. When you sunset your latest flagship because the economics don’t close, that’s not a moon shot trajectory, it’s weight shaving on a treehouse.

Stanford’s 2025 AI-Index shows that model to model spreads on MMLU, HumanEval, and GSM8K have collapsed to low single digits, performance curves are flattening exactly where compute curves are exploding. A fresh MIT-CSAIL paper modelling "Bayes slowdown" makes the same point mathematically: every extra order of magnitude of FLOPs is buying less accuracy than the one before.[1]

A survey published last week[2] catalogs the 2025 state of RLHF/RLAIF: reward hacking, preference data scarcity, and training instability remain open problems, just mitigated by ever heavier regularisation and bigger human in the loop funnels. If our alignment patch still needs a small army of labelers and a KL muzzle to keep the model from self lobotomising calling it "solved" feels optimistic.

Scale, fancy sampling tricks, and patched up RL got us to the leafy top so chatbots that can code and debate decently. But the same reports above show the branches bending under compute cost, data saturation, and alignment tax. Until we swap out the propulsion system so new architectures, richer memory, or learning paradigms that add information instead of reweighting it we’re in danger of planting a flag on a treetop and mistaking it for Mare Tranquillitatis.

Happy to climb higher together friend but I’m still packing a parachute, not a space suit.

1. https://arxiv.org/html/2507.07931v1

2. https://arxiv.org/html/2507.04136v1

Re: ChatGPT agent: bridging research and action

#385

Earlier quoted context omitted.

This is why an on device browser is coming. It'll let the AI platforms get around any other platform blocks by hijacking the consumer's browser. And it makes total sense, but hopefully everyone else has done the game theory at least a step or two beyond that.

You mean like calaude code's integration with play right ?

No, because playwright can be detected pretty easily and blocked. It needs to be (and will be) using the same browser that you regularly browse with.

Re: ChatGPT agent: bridging research and action

#386

The security risks with this sound scary. Let's say you give it access to your email and calendar. Now it knows all of your deepest secrets. The linked article acknowledges that prompt injection is a risk for the agent: > Prompt injections are attempts by third parties to manipulate its behavior through malicious instructions that ChatGPT agent may encounter on the web while completing a task. For example, a maliciou…

The asking for permission thing is irrelevant. People are using this tool to get the friction in their life to near zero, I bet my job that everyone will just turn on auto accept and go for a walk with their dog.

Re: ChatGPT agent: bridging research and action

#387
Nice action plan on combining Operator and Deep Research.

One thing which stood out to me in a thought-provoking way, is that example of stickers [created first and then] being ordered (obviously: pending ordering confirmation from the user) from StickerSpark (JFYI: This is a fictional company made up in this OpenAI launch post), whereby as mentioned that ChatGPT agent has "its own computer". Thus, if OpenAI is logging into its own account on StickerSpark, then what would be StickerSpark's "normal" user-base like that of any other company's user-base of 1 user per actual person will shift to StickerSpark having a few large users via agents through OpenAI, Anthropic, Google, etc. and a medium-long tail of regular individual users. This exactly reminds of how through pervasive index fund investing that index fund houses such as BlackRock and Vanguard directly own large stakes in many S&P500 companies such that they can sway voting power [1]. Thus, with ChatGPT agent that the fundamental-regular-interaction that we assume with websites like StickerSpark would stand to alter whereby the agents would be business-facing and would have more influence on the website's features (or the Agent due to its innate intelligence will directly find another website for where features match up).

[1] https://manhattan.institute/article/index-funds-have-too-muc...

Re: ChatGPT agent: bridging research and action

#388
post #225

Earlier quoted context omitted.

The proper use of these systems is to treat them like an intern or new grad hire. You can give them the work that none of the mid-tier or senior people want to do, thereby speeding up the team. But you will have to review their work thoroughly because there is a good chance they have no idea what they are actually doing. If you give them mission-critical work that demands accuracy or just let them have free rein with…

I’ve never experienced an intern who was remotely as mediocre and incapable of growth as an LLM.

I had an intern who didn’t shower. We had to have discussions about body odor in an office. AI/LLM’s are an improvement in that regard. They also do better work than that kid did. At least he had rich parents.

Re: ChatGPT agent: bridging research and action

#389
post #148

Earlier quoted context omitted.

> said we'd have self driving cars "in a few years" back in 2015 And they wouldn't have been too far off! Waymo became L4 self-driving in 2021, and has been transporting people in the SF Bay Area without human supervision ever since. There are still barriers — cost, policies, trust — but the technology certainly is here.

People were saying we would all be getting in our cars and taking a nap on our morning commute. We are clearly still a pretty long ways off from self-driving being as ubiquitous as it was claimed it would be.

And other people were a lot more moderate but still assumed we'd get self-driving soon, with caveats, and were bang on the money.

So it's not as ubiquitous as the most optimistic estimates suggested. We're still at a stage where the tech is sufficiently advanced that seeing them replace a large proportion of human taxi services now seems likely to have been reduced to a scaling / rollout problem rather than primarily a technology problem, and that's a gigantic leap.

Re: ChatGPT agent: bridging research and action

#390
post #336

This solves a big issue for existing CLI agents, which is session persistence for users working from their own machines. With claude code, you usually start it from your own local terminal. Then you have access to all the code bases and other context you need and can provide that to the AI. But when you shut your laptop, or have network availability changes the show stops. I've solved this somewhat on MacOS using the…

What tasks are you running that take more than a few minutes without intervention?

When using spec writter and sub-tasking tools like TaskMaster, Kiro, etc. I've experienced Claude Code to take 30-60+ minutes for a more complex feature
Post reply on HN