Live data from Hacker News

Operator research preview

openai.com

171–180 of 448 posts

Re: Operator research preview

#171

Overall, Operator seems the same as Claude's Computer Use demo from a few months ago, including architecture requiring user to launch a VM, and a tendency to be incorrect: https://news.ycombinator.com/item?id=41914989 Notably, Claude's Computer Use implementation made few waves in the AI Agent industry since that announcement despite the hype.

Correction on "including architecture requiring user to launch a VM": apparently OpenAI uses a cloud hosted VM that's shown to the user. While that's much more user friendly, it opens up different issues around security/privacy.

Re: Operator research preview

#172

> and even creating memes. important work. glad to hear they're investing $500B in this space instead of stuff like, I don't know, making the planet livable for our grandkids

"Operator, I need to purchase 78,000 widgets for my company. Please find the best deal among suppliers who ship using carriers and ports who meet or exceed US EPA guidelines. Please ensure at least 50% of the product is sourced from post-consumer waste, and order your responses by price per unit."

Re: Operator research preview

#173

What is fascinating about this announcement is if you look into future after considerable improvements in product and the model, we will be just chatting with ChatGPT to book dinner tables, flights, buy groceries and do all sort of mundane and hugely boring things we do on the web, just by talking to the agents. I'd definitely love that.

I don't. Chat interface sucks; for most of these things, a more direct interface could be much more ergonomic, and easier to operate and integrate. The only reason we don't have those interfaces is because neither restaurants, nor airlines, nor online stores, nor any other businesses actually want us to have them. To a business, the user interface isn't there to help the user achieve their goals - it's a platform for…

I also do not like Chat interface. What I meant by above comment was actually talking and having natural conversations with Operator agent while driving car or just going for a walk or whenever and wherever something comes to my mind which requires me to go to browser and fill out forms etc. That would get us closer to using chatGPT as a universal AI agent to get those things done. (This is what Siri was supposed to be one day when Steve Jobs introduced it on that stage but unfortunately that day never arrived.)

Re: Operator research preview

#175
post #71

Earlier quoted context omitted.

> I'm not convinced this is the path forward for computers either though. With this approach they'll have to contend with the agent running into all the anti-bot measures that sites have implemented to deal with abuse. CAPTCHAs, flagging or blocking datacenter IP addresses, etc. Maybe deals could be struck to allow agents to be whitelisted, but that assumes the agents won't also be used for abuse. If you could get Ch…

The solution is simple, and it's what's already done with search by proprietary LLMs: reasoning happens on the LLM vendor's servers, tool use happens client-side . Whether for search or "computer use", the websites will register activity coming from the user's machine, as it should be, because LLMs act as User Agents here . Of course, already with LLM-powered search we see growing number of people doing the selfish/i…

[deleted]

Re: Operator research preview

#176

What is fascinating about this announcement is if you look into future after considerable improvements in product and the model, we will be just chatting with ChatGPT to book dinner tables, flights, buy groceries and do all sort of mundane and hugely boring things we do on the web, just by talking to the agents. I'd definitely love that.

I would really love for Apple Knowledge Navigator to be real: https://www.youtube.com/watch?v=umJsITGzXd0

and I'm surprised that people don't bring this visualisation up more often.

Re: Operator research preview

#177

What is fascinating about this announcement is if you look into future after considerable improvements in product and the model, we will be just chatting with ChatGPT to book dinner tables, flights, buy groceries and do all sort of mundane and hugely boring things we do on the web, just by talking to the agents. I'd definitely love that.

Are our attention spans so shot that we consider booking a reservation at a restaurant or buying groceries "hugely boring"? And do we value convenience so much that we're willing to sacrifice a huge breadth of options for whatever sponsor du jour OpenAI wants to serve us just to save less than 10 minutes?

And would this company spend billions of dollars for this infinitesimally small increase in convenience? No, of course not; you are not the real customer here. Consider reading between the lines and thinking about what you are sacrificing just for the sake of minor convenience.

Re: Operator research preview

#178

From the slide deck on the livestream: "[Operator safety risks and mitigations] Harmful tasks: User is misaligned" Looking forward to seeing some more of the examples for when openai considers their users as "misaligned", whatever that actually even means anymore.

I assume here it means complying with requests that could harm other people. It's pretty common for businesses to tell their employees not to assist customers doing bad things, so not surprised to see AIs trained to not to assist customers doing bad things. Examples: - "operator, please sign up for 100 fake Reddit accounts and have them regularly make posts praising product X." - "operator, please order the component…

I appreciate that they all say please.

Re: Operator research preview

#179
post #31

I don't know why, but the approach where "agents" accomplish things by using a mouse and keyboard and looking at pixels always seemed off to me. I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes. Ima…

You could make a similar argument for self-driving cars. We would have got there quicker if the roads were built from the ground up for automation. You can try to get the world on board to change how they do roads. Or make the computers adapt to any kind of road.

Re: Operator research preview

#180

I sometimes wonder if Rabbit and their LAM ( https://www.rabbit.tech/lam-playground ) were just a year too early to market.

The issue with rabbit is that their flagship product was a poorly disguised android device that tapped into vanilla ChatGPT, when it was marketed as "the thing that will replace smartphones".
Post reply on HN