Live data from Hacker News

ChatGPT agent: bridging research and action

openai.com

51–60 of 508 posts

Re: ChatGPT agent: bridging research and action

#52
post #38
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

> It feels like either finding that 2% that's off (or dealing with 2% error) will be the time consuming part in a lot of cases. The last '2%' (and in some benchmarks 20%) could cost as much as $100B+ more to make it perfect consistently without error. This requirement does not apply to generating art. But for agentic tasks, errors at worst being 20% or at best being 2% for an agent may be unacceptable for mistakes. A…

[deleted]

Re: ChatGPT agent: bridging research and action

#53
post #21
post #5

[flagged]

By 2030 Europe will be known for croissants and colossal brains.

The European livestyle isn't god given and has to be paid for. It's a luxury and I'm still puzzled that people don't get that we can't afford it without an economy.

Re: ChatGPT agent: bridging research and action

#54
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

Lol the music and presentation made it sound like that guy was going to talk about something deep and emotional not spreadsheets and expense reports.

Re: ChatGPT agent: bridging research and action

#55
Whilst we have seen other implementations of this (providing a VPS to an LLM), this does have a distinct edge others in the way it presents itself. The UI shown, with the text overlay, readable mouse and tailored UI components looks very visually appealing and lends itself well to keeping users informed on what is happening and why at every stage. I have to tip my head to OpenAIs UI team here, this is a really great implementation and I always get rather fascinated whenever I see LLMs being implemented in a visually informative and distinctive manner that goes beyond established metaphors.

Comparing it to the Claude+XFCE solutions we have seen by some providers, I see little in the way of a functional edge OpenAI has at the moment, but the presentation is so well thought out that I can see this being more pleasant to use purely due to that. Many times with the mentioned implementations, I struggled with readability. Not afraid to admit that I may borrow some of their ideas for a personal project.

Re: ChatGPT agent: bridging research and action

#56

It's smart that they're pivoting to using the user's computer directly - managing passwords, access control and not getting blocked was the biggest issue with their operator release. Especially as the web becomes more and more locked down. > ChatGPT agent's output is comparable to or better than that of humans in roughly half the cases across a range of task completion times, while significantly outperforming o3 and…

Doesn't the very first line say the opposite?

"ChatGPT can now do work for you using its own computer"

Re: ChatGPT agent: bridging research and action

#59
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

I think this is my favorite part of the LLM hype train: the butterfly effect of dependence on an undependable stochastic system propagates errors up the chain until the whole system is worthless.

"I think it got 98% of the information correct..." how do you know how much is correct without doing the whole thing properly yourself?

The two options are:

- Do the whole thing yourself to validate

- Skim 40% of it, 'seems right to me', accept the slop and send it off to the next sucker to plug into his agent.

I think the funny part is that humans are not exempt from similar mistakes, but a human making those mistakes again and again would get fired. Meanwhile an agent that you accept to get only 98% of things right is meeting expectations.

Re: ChatGPT agent: bridging research and action

#60
post #48
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

I’ve worked at places that sre run on spreadsheets. You’d be amazed at how often they’re wrong IME

It takes my boss seven hours to create that spreadsheet, and another eight to render a graph.
Post reply on HN