Operator research preview
91–100 of 448 posts
Re: Operator research preview
#92Earlier quoted context omitted.
[flagged]
Then why not say that in the first place? It's not very charitable to get upset with someone for correcting an incorrect number you yourself stated. Also, what's up with the faux-stutter?
Do you understand how much money a hundred billion dollars is?
Re: Operator research preview
#93Wonder what’s changed recently..
Re: Operator research preview
#94They also just announced o3-mini will be on free tier for chatGPT as well.
Re: Operator research preview
#95Re: Operator research preview
#96I don't know why, but the approach where "agents" accomplish things by using a mouse and keyboard and looking at pixels always seemed off to me. I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes. Ima…
> but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf I sincerely hope it's not the future we're heading to (but it might be inevitable, sadly). If it becomes a popular trend, developers will start making "AI-first" apps that you have to use AI to interact with to get the full functionality. See also: mobile first.
The developer's incentive is to control the experience for a mix of the users' ends and the developer's ends. Functionality being what users want and monetization being what developers want. Devs don't expose APIs for the same reason why hackers want them - it commodifies the service.
An AI-first app only makes sense if the developer controls the AI and is developing the app to sell AI subscriptions. An independent AI company has no incentive to support the dev's monetization and every incentive to subvert it in favor of their own.
(EDIT: This is also why AI agents will "use" mice and keyboards. The agent provider needs the app or service to think they're interacting with the actual human user instead of a bot, or else they'll get blocked.)
Re: Operator research preview
#97I don't know why, but the approach where "agents" accomplish things by using a mouse and keyboard and looking at pixels always seemed off to me. I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes. Ima…
> I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf OS specific, but Apple has the Scripting Support API [0] and Shortcut API for their app. Works great. [0]: https://developer.apple.com/documentation/foundation/scripti...
Re: Operator research preview
#98Earlier quoted context omitted.
Big jumps in benchmarks from Claude's Computer Use though. 87% vs 56% on Webvoyager 58.1% vs 36.2% on WebArena 38.1% vs 22% on OsWorld These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)
OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-use agents (which can use the DOM, browser-specific APIs, and so on.) The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with…
Re: Operator research preview
#99I dream of a world where I can specify annoying things to me and build a perfect search for any house, that understands how I think about money, how I think about my family, and what I love and really extends how I interact with the world.
Re: Operator research preview
#100We already have this https://github.com/browser-use/browser-use