Live data from Hacker News

ChatGPT agent: bridging research and action

openai.com

321–330 of 508 posts

Re: ChatGPT agent: bridging research and action

#321

Earlier quoted context omitted.

It’s past the hype curve and into the trough of disillusionment. Over the next 5,10,15 years (who can say?) the tech will mature out of the trough into general adoption. GenAI is the exciting new tech currently riding the initial hype spike. This will die down into the trough of disillusionment as well, probably sometime next year. Like self-driving, people will continue to innovate in the space and the tech will be…

The Gartner hype cycle assumes a single fundamental technical breakthrough, and describes the process of the market figuring out what it is and isn't good for. This isn't straightforwardly applicable to LLMs because the question of what they're good for is a moving target; the foundation models are actually getting more capable every few months, which wasn't true of cryptocurrency or self-driving cars. At least some…

The Gartner hype cycle is complete nonsense, it's just a completely fabricated way to view the world that helps sell Gartner's research products. It may, at times, make "intuitive sense", but so does astrology.

The hype cycle has no mathematical basis whatsoever. It's marketing gimmick. It's only value in my life has been to quickly identify people that don't really understand models or larger trends in technology.

I continue to be, but on introspection probably shouldn't be, surprised that people on HN treat is as some kind of gospel. The only people who should respected are other people in the research marketing space as the perfect example of how to dupe people into paying for your "insights".

Re: ChatGPT agent: bridging research and action

#322
post #213

Earlier quoted context omitted.

In general most of the previous AI "breakthrough" in the last decade were backed by proper scientific research and ideas: - AlphaGo/AlphaZero (MCTS) - OpenAI Five (PPO) - GPT 1/2/3 (Transformers) - Dall-e 1/2, Stable Diffusion (CLIP, Diffusion) - ChatGPT (RLHF) - SORA (Diffusion Transformers) "Agents" is a marketing term and isn't backed by anything. There is little data available, so it's hard to have generally capa…

Yep. Agents are only powered by clever use of training data, nothing more. There hasn't been a real breakthrough in a long time.

"Long time" as in, 7 months since o1 and reasoning models were released? That was a pretty big breakthrough.

Re: ChatGPT agent: bridging research and action

#323
post #308

Earlier quoted context omitted.

“The quip about 98% correct should be a red flag for anyone familiar with spreadsheets” I disagree. Receiving a spreadsheet from a junior means I need to check it. If this gives me infinite additional juniors I’m good. It’s this popular pattern of HN comments - expect AI to behave deterministically correct - while the whole world operates on stochastically correct all the time…

In my experience the value of junior contributors is that they will one day become senior contributors. Their work as juniors tends to require so much oversight and coaching from seniors that they are a net negative on forward progress in the short term, but the payoff is huge in the long term.

Exactly this

And it should go without saying that LLMs do not have the same investment/value tradeoff. Whether or not they contribute like a senior or junior seems entirely up to luck

Prompt skill is flaky and unreliable to ensure good output from LLMs

Re: ChatGPT agent: bridging research and action

#324
post #14

It's very hard for me to imagine the current level of agents serving a useful purpose in my personal life. If I ask this to plan a date night with my wife this weekend, it needs to consult my calendar to pick the best night, pick a bar and restaurant we like (how would it know?), book a babysitter (can it learn who we use and text them on my behalf?), etc. This is a lot of stuff it has to get right, and it requires a…

This problem particularly interests me. One of my favorite use cases for these tools is travel where I can get recommendations for what to do and see without SEO content. This workflow is nice because you can ask specific questions about a destination (e.g., historical significance, benchmark against other places). ChatGPT struggles with: - my current location - the current time - the weather - booking attractions an…

The best resource I've found for travel is travel forums. Asking any AI so far it mostly feeds me the same SEO content, but packaged up a bit nicer.

Re: ChatGPT agent: bridging research and action

#325
post #238

Earlier quoted context omitted.

It's "here" if you live in a handful of cities around the world, and travel within specific areas in those cities. Getting this tech deployed globally will take another decade or two, optimistically speaking.

Given how well it seems to be going in those specific areas, it seems like it's more of a regulatory issue than a technological one.

Maybe, but it's also going to be a financial issue eventually too

My city had Car2Go for a couple of years, but it's gone now. They had to pull out of the region because it wasn't making them enough money

I expect Waymo and any other sort of vehicle ridesharing thing will have the same problem in many places

Re: ChatGPT agent: bridging research and action

#326
post #30

The "spreadsheet" example video is kind of funny: guy talks about how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... I just needed to copy / paste a few things. If it can do 90 - 95% of the…

> how it normally takes him 4 to 8 hours to put together complicated, data-heavy reports. Now he fires off an agent request, goes to walk his dog, and comes back to a downloadable spreadsheet of dense data, which he pulls up and says "I think it got 98% of the information correct... This is where the AI hype bites people. A great use of AI in this situation would be to automate the collection and checking of data. Se…

98% sure each commit doesn’t corrupt the database, regress a customer feature, open a security vulnerability. 50 commits later … (which is like, one day for an agentic workflow)

Re: ChatGPT agent: bridging research and action

#327

Earlier quoted context omitted.

Of course, Pareto principle is at work here. In an adjacent field, self-driving, they are working on the last "20%" for almost a decade now. It feels kind of odd that almost no one is talking about self-driving now, compared to how hot of a topic it used to be, with a lot of deep, moral, almost philosophical discussions.

> It feels kind of odd that almost no one is talking about self-driving now, compared to how hot of a topic it used to be Probably because it's just here now? More people take Waymo than Lyft each day in SF.

Yeah where they have every inch of SF mapped, and then still have human interventions. We were promised no more human drivers like 5-7 years ago at this point.

Re: ChatGPT agent: bridging research and action

#328
post #70

I've been using OpenAI operator for some time - but more and more websites are blocking it, such as LinkedIn and Amazon. That's two key use-cases gone (applying to jobs and online shopping). Operator is pretty low-key, but once Agent starts getting popular, more sites will block it. They'll need to allow a proxy configuration or something like that.

If people will actually pay for stuff (food, clothing, flights, whatever) through this agent or operator, I see no reason Amazon etc would continue to block them.

Possibly in part because bots will not fall for the same tricks as humans (recommended items, as well as other things which amazon does to try and get the most money possible)

Re: ChatGPT agent: bridging research and action

#329

Earlier quoted context omitted.

Yep. Agents are only powered by clever use of training data, nothing more. There hasn't been a real breakthrough in a long time.

"Long time" as in, 7 months since o1 and reasoning models were released? That was a pretty big breakthrough.

In the context of our conversation and what OP wrote, there has been no breakthrough since around 2018. What you're seeing is the harvesting of all low-hanging fruit from a tree that was discovered years ago. But fruit is almost gone. All top models perform at almost the same level. All the "agents" and "reasoning models" are just products of training data.

I wrote more about it here:

https://news.ycombinator.com/item?id=44426993

You may also be interested in this article, that goes into details even more:

https://blog.jxmo.io/p/there-are-no-new-ideas-in-ai-only

Re: ChatGPT agent: bridging research and action

#330

Earlier quoted context omitted.

Of course, Pareto principle is at work here. In an adjacent field, self-driving, they are working on the last "20%" for almost a decade now. It feels kind of odd that almost no one is talking about self-driving now, compared to how hot of a topic it used to be, with a lot of deep, moral, almost philosophical discussions.

The critics of the current AI buzz certainly have been drawing comparisons to self driving cars as LLMs inch along with their logarithmic curve of improvement that's been clear since the GPT-2 days. Whenever someone tells me how these models are going to make white collar professions obsolete in five years, I remind them that the people making these predictions 1) said we'd have self driving cars "in a few years" bac…

How profound. No one has ever posted that exact same thought before on here. Thank you.
Post reply on HN