Live data from Hacker News

ChatGPT agent: bridging research and action

openai.com

411–420 of 508 posts

Re: ChatGPT agent: bridging research and action

#411
post #336

Earlier quoted context omitted.

What tasks are you running that take more than a few minutes without intervention?

When using spec writter and sub-tasking tools like TaskMaster, Kiro, etc. I've experienced Claude Code to take 30-60+ minutes for a more complex feature

Could you explain in more detail how you’re using these tools?

Re: ChatGPT agent: bridging research and action

#412
The demo will be great and it will not be accurate enough or trustworthy enough to touch, however many people will start automating their jobs with it and producing absolute crap on fairly important things. We are moving from people just making things (post-truth) to the actual information all being corrupted (post-correct? there's got to be a better shorthand for this).

Re: ChatGPT agent: bridging research and action

#413
Yeah but not opensource. The community is pissed off about it, Openai should contribute back to the community. Reddit blow up with angry users about the delays and missed promises of Openai release models. The fact they thay are affraid to be humiliated from the competition like "llama4 did" is not an excuse, in fact should be the motivation.

Re: ChatGPT agent: bridging research and action

#414

Earlier quoted context omitted.

In the system I'm building the main agent doesn't have access to tools and must call scoped down subagents who have one or two tools at most and always in the same category (so no mixed fetch and calendar tools). They must also return structured data to the main agent. I think that kind of isolation is necessary even though it's a bit more costly. However since the subagents have simple tasks I can use super cheap mo…

What isolation is there? If a compromised sub agent returns data that gets inserted into the main agents context (structured or not) then the end result is the same as if the main agent was directly interacting with the compromising resource is it not?

Exactly. You can't both give the model access AND enforce security. You CAN convince yourself you've done it though. You see it all the time, including in this thread.

Re: ChatGPT agent: bridging research and action

#415
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

My favorite part is people taking the 98% number to heart as if there's any basis to it whatsoever and isn't just a number they pulled out of their ass in this marketing material made by an AI company trying to sell you their AI product. In my experience it's more like a 70% for dead simple stuff, and dramatically lower for anything moderately complex.

And why 98%? Why not 99% right? Or 99.9% right? I know they can't outright say 100% because everyone knows that's a blatant lie, but we're okay with them bullshitting about the 98% number here?

Also there's no universe in which this guy gets to walk his dog while his little pet AI does his work for him, instead his boss is going to hound him into doing quadruple the work because he's now so "efficient" that he's finishing his spreadsheet in an hour instead of 8 or whatever. That, or he just gets fired and the underpaid (or maybe not even paid) intern shoots off the same prompt to the magic little AI and does the same shoddy work instead of him. The latter is definitely what the C-suite is aiming for with this tech anyway.

Re: ChatGPT agent: bridging research and action

#416
post #225

Earlier quoted context omitted.

The proper use of these systems is to treat them like an intern or new grad hire. You can give them the work that none of the mid-tier or senior people want to do, thereby speeding up the team. But you will have to review their work thoroughly because there is a good chance they have no idea what they are actually doing. If you give them mission-critical work that demands accuracy or just let them have free rein with…

What a awful way to think about internship. The goal is to help people grow, so they can achieve things they would not have been able to deal with before gaining that additional experience. This might include boring dirty work, yes. But that means they thus prove they can overcome such a struggle, and so more experienced people should be expected to also be able to go though it - if there is no obvious more pleasant…

I agree that it sounds harsh. But I worked for a company that hired interns and this was the way that managers talked about them- as cheap, unreliable labor. I once spoke with an intern hoping that they could help with a real task: using TensorFlow (it was a long time ago) to help analyze our work process history, but the company ended up putting them on menial IT tasks and they checked out mentally.

Re: ChatGPT agent: bridging research and action

#417
post #107

Earlier quoted context omitted.

THIS is the main problem. I was listening the whole time for them to announce a way to run it locally or at least proxy through your local devices. Alas the Deepseek R1 distillation experience they went through (a bit like when Steve Jobs was fuming at Google for getting Android to market so quickly) made them wary of showing to many intermediate results, tricks etc. Even in the very beginning Operator v1 was unable…

This is why an on device browser is coming. It'll let the AI platforms get around any other platform blocks by hijacking the consumer's browser. And it makes total sense, but hopefully everyone else has done the game theory at least a step or two beyond that.

Perplexity already launched it.

Re: ChatGPT agent: bridging research and action

#420
post #65

Very slightly impressed by their emphasis on the gigantic (my word, not theirs) risk of giving the thing access to real creds and sensitive info.

Their target demographic was already born into the matrix and they don't even know it's there; it will hardly be a problem for them.
Post reply on HN