Live data from Hacker News

Operator research preview

openai.com

121–130 of 448 posts

Re: Operator research preview

#121
I strongly believe we need to use Open APIs for agents. OpenAPI is the perfect specification standard that would allow for an open world and an open internet for agents.

When OpenAI first came out with their first version of GPTs, it was all based on open APIs.

Now they are moving away from it more and more. This means they want to control the market because they don't want to base it on an open standard.

It's such a shame!

Re: Operator research preview

#122

Earlier quoted context omitted.

Big jumps in benchmarks from Claude's Computer Use though. 87% vs 56% on Webvoyager 58.1% vs 36.2% on WebArena 38.1% vs 22% on OsWorld These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)

OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-use agents (which can use the DOM, browser-specific APIs, and so on.) The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with…

> OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents.

No. It's not matching them, it's clearly exceeding them. The previous post provided the numbers.

Re: Operator research preview

#123

Waiting for the "OpenAI has no moat" crowd to chime in while they keep releasing new features and dominating market share. (And yeah, they just got half a trillion ). Edit: Downvote all you want, reality won't change. Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing. What about "AI will never replace Hollywood actors". What about that time…

The announcement of investing $500B with the proposed benefit of creating $100K jobs - to my amazement did not produce any commentary that I came across raising questions about the ROI of spending $5M per job created. I mean it’s all right there in the announcement! For instance, the American Recovery and Reinvestment Act (ARRA) of 2009, which allocated approximately $787 billion, was estimated to have created or sav…

They also gave no indication of what types of jobs were going to be created. It's pretty hot-air.

If the goal is to create data centers for more AI training, you can rest assured that depends on creating as few jobs as possible in order to keep labor costs down and have more to spend on hardware and energy.

Re: Operator research preview

#124

Waiting for the "OpenAI has no moat" crowd to chime in while they keep releasing new features and dominating market share. (And yeah, they just got half a trillion ). Edit: Downvote all you want, reality won't change. Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing. What about "AI will never replace Hollywood actors". What about that time…

The announcement of investing $500B with the proposed benefit of creating $100K jobs - to my amazement did not produce any commentary that I came across raising questions about the ROI of spending $5M per job created. I mean it’s all right there in the announcement! For instance, the American Recovery and Reinvestment Act (ARRA) of 2009, which allocated approximately $787 billion, was estimated to have created or sav…

None of the $500B comes from the government, so the cost is $0 of government spending per job.

Re: Operator research preview

#125

Earlier quoted context omitted.

OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-use agents (which can use the DOM, browser-specific APIs, and so on.) The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with…

> OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. No. It's not matching them, it's clearly exceeding them. The previous post provided the numbers.

[deleted]

Re: Operator research preview

#126

Earlier quoted context omitted.

OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-use agents (which can use the DOM, browser-specific APIs, and so on.) The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with…

> OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. No. It's not matching them, it's clearly exceeding them. The previous post provided the numbers.

Those numbers are not the full story. Note that GP specifically says: "Big jumps in benchmarks from _Claude's Computer Use_ though." Claude Computer Use was not SOTA for browser tasks at the time of its release (and is still not.)

In WebArena, Operator does 58.1%. Previous SOTA for browser-use agents is 57.1%. In WebVoyager, Operator does 87.0%. Previous SOTA for browser-use agents is the exact same.

See here for details: https://openai.com/index/computer-using-agent/

Re: Operator research preview

#127

I strongly believe we need to use Open APIs for agents. OpenAPI is the perfect specification standard that would allow for an open world and an open internet for agents. When OpenAI first came out with their first version of GPTs, it was all based on open APIs. Now they are moving away from it more and more. This means they want to control the market because they don't want to base it on an open standard. It's such a…

Models will eventually be interface agnostic and they will cover all interfaces that are commonly used by individuals and organizations. It won't matter whether you have a nicely documented public API, a traditional website, or a phone interface to customer support.

Re: Operator research preview

#128
Curious how long this paradigm (computers using human interfaces) will last for P95 tasks.

If the machines are smart enough, shouldn’t they be able to build better interfaces to existing software?

With that aside, it seems like there are two things at play in this demo:

1. Pixel-tuned GPT-4o

2. “Agent” in prod (supervisor loop + operator loop)

Will be interesting to see if they open those up as separate tools in the future, or if they let this fall to the wayside like GPTs, Dalle, etc.

Re: Operator research preview

#129

Overall, Operator seems the same as Claude's Computer Use demo from a few months ago, including architecture requiring user to launch a VM, and a tendency to be incorrect: https://news.ycombinator.com/item?id=41914989 Notably, Claude's Computer Use implementation made few waves in the AI Agent industry since that announcement despite the hype.

Big jumps in benchmarks from Claude's Computer Use though. 87% vs 56% on Webvoyager 58.1% vs 36.2% on WebArena 38.1% vs 22% on OsWorld These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)

Gemini is 90.5% in Webvoyager[1] compared to 87% for OpenAI.

[1]: https://deepmind.google/technologies/project-mariner/

Re: Operator research preview

#130
post #117

This space is moving fast. You can now run a local open model to control your browser or entire computer: https://github.com/bytedance/UI-TARS

I saw this earlier. benchmarks are impressive!

Did OpenAI release anything beside this product? Any benchmarks at least to compare?

It feels like OpenAI is betting on the fact that they have a nice UI?!

Post reply on HN