Live data from Hacker News

ChatGPT agent: bridging research and action

openai.com

211–220 of 508 posts

Re: ChatGPT agent: bridging research and action

#211
post #176

Predicted by the AI 2027 team in early April: > Mid 2025: Stumbling Agents The world sees its first glimpse of AI agents. Advertisements for computer-using agents emphasize the term “personal assistant”: you can prompt them with tasks like “order me a burrito on DoorDash” or “open my budget spreadsheet and sum this month’s expenses.” They will check in with you as needed: for example, to ask you to confirm purchases.…

It was common knowledge that big corps were working on agent-type products when that report was written. Hardly much of a prediction, let alone any sort of technical revolution.

Re: ChatGPT agent: bridging research and action

#212
post #5

[flagged]

Hardly. Is Apple a doomed company because they are chronically late to ~everything bleeding edge?

*glances at AI, VR, mini phones, smart cars, multi-wireless charging, home automation, voice assistants, streaming services, set-top boxes, digital backup software, broadband routers, server hardware, server software and 12" laptops in rapid succession*

Maybe(!?!)

Re: ChatGPT agent: bridging research and action

#213

I'm not so optimistic as someone that works on agents for businesses and creating tools for it. The leap from low 90s to 99% is classic last mile problem for LLM agents. The more generic and spread an agent is (can-do-it-all) the more likely it will fail and disappoint. Can't help but feel many are optimizing happy paths in their demos and hiding the true reality. Doesn't mean there isn't a place for agents but rathe…

In general most of the previous AI "breakthrough" in the last decade were backed by proper scientific research and ideas:

- AlphaGo/AlphaZero (MCTS)

- OpenAI Five (PPO)

- GPT 1/2/3 (Transformers)

- Dall-e 1/2, Stable Diffusion (CLIP, Diffusion)

- ChatGPT (RLHF)

- SORA (Diffusion Transformers)

"Agents" is a marketing term and isn't backed by anything. There is little data available, so it's hard to have generally capable agents in the sense that LLMs are generally capable

Re: ChatGPT agent: bridging research and action

#215

Earlier quoted context omitted.

There are in fact lots of tasks people complete immediately at 99.99% success rate at first iteration or 99.999% after self and peer checking work Perhaps importantly checking is a continual process and errors are identified as they are made and corrected whilst in context instead of being identified later by someone completely devoid of any context a task humans are notably bad at. Lastly it's important to note the…

> There are in fact lots of tasks people complete immediately at 99.99% success rate at first iteration or 99.999% after self and peer checking work This is so absurd that I wonder if you're telling? Humans don't even have a 99.99% success rate in breathing, let alone any cognitive tasks.

> Humans don't even have a 99.99% success rate in breathing

Will you please elaborate a little on this?

Re: ChatGPT agent: bridging research and action

#216
post #213

I'm not so optimistic as someone that works on agents for businesses and creating tools for it. The leap from low 90s to 99% is classic last mile problem for LLM agents. The more generic and spread an agent is (can-do-it-all) the more likely it will fail and disappoint. Can't help but feel many are optimizing happy paths in their demos and hiding the true reality. Doesn't mean there isn't a place for agents but rathe…

In general most of the previous AI "breakthrough" in the last decade were backed by proper scientific research and ideas: - AlphaGo/AlphaZero (MCTS) - OpenAI Five (PPO) - GPT 1/2/3 (Transformers) - Dall-e 1/2, Stable Diffusion (CLIP, Diffusion) - ChatGPT (RLHF) - SORA (Diffusion Transformers) "Agents" is a marketing term and isn't backed by anything. There is little data available, so it's hard to have generally capa…

My personal framing of "Agents" is that they're more like software robots than they are an atomic unit of technology. Composed of many individual breakthroughs, but ultimately a feat of design and engineering to make them useful for a particular task.

Re: ChatGPT agent: bridging research and action

#218
And I'm still waiting for the simple feature – the ability to edit documents in projects.

I use projects for working on different documents - articles, research, scripts, etc. And would absolutely love to write it paragraph after paragraph with the help of ChatGPT for phrasing and using the project knowledge. Or using voice mode - i.e. on a walk "Hey, where did we finish that document - let's continue. Read the last two paragraphs to me... Okay, I want to elaborate on ...".

I feel like AI agents for coding are advancing at a breakneck speed, but assistance in writing is still limited to copy-pasting.

Re: ChatGPT agent: bridging research and action

#219
post #198

Earlier quoted context omitted.

Hardly. Is Apple a doomed company because they are chronically late to ~everything bleeding edge?

Apple products are leading edge. Imagine if they waited until Samsung makes the perfect phone , then copy it. We re talking about european tech businesses being left behind, locked in a basement.

So you have a positive opinion when Apple does things after others, but Europe having a slower, cautious approach is treated as negative for you?

What is your preference for Europe, complete floodgates open and never ending lawsuits over IP theft like we have in the USA currently over AI?

The US is not the example of what’s working, it’s merely a demonstration of what is possible when you have limited, provoked regulation.

Re: ChatGPT agent: bridging research and action

#220

One the one hand this is super cool and maybe very beneficial, something I definitely want to try out. On the other, LLMs always make mistakes, and when it's this deeply integrated into other system I wonder how severe these mistakes will be, since they are bound to happen.

Based on the live stream, so does OpenAI.

But of course humans makes a multitude of mistakes too.

Post reply on HN