Live data from Hacker News

Operator research preview

openai.com

61–70 of 448 posts

Re: Operator research preview

#61
post #6

As usual this is quite underwhelming. All this hype and it appears that this was a rushed last minute demo to show something that is hardly ready. Ever since GPTs, "Operator" looks quite frankly gimmicky.

I agree with this one. But you have to start somewhere. I think in the next several things, websites will be built for agents and not people. So it'll only get better and smarter.

Will they? What incentive is there if people haven't started using agents yet?

We already have a way to build websites for machines: it's called APIs. And frankly, I think that's a better answer for "hooking LLM into website" -- the things which make APIs hard for humans (discomfort, inconvenience, low discoverability, technical complexity) aren't really problems for LLMs.

Re: Operator research preview

#63

Earlier quoted context omitted.

I thought it was a general fund for ai related stuff. Was this all for a single company?

It's for OpenAI.

I'm looking at the reuters article: https://www.reuters.com/technology/artificial-intelligence/t...

Unless it's misrepresented, this looks like an earmarked VC fund

Re: Operator research preview

#64
post #31

I don't know why, but the approach where "agents" accomplish things by using a mouse and keyboard and looking at pixels always seemed off to me. I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes. Ima…

probably more difficult to get app developers on board.

You answered your own question. You have to build the ecosystem if you want to have the facilities your comment outlines.

Whereas the facilities are already in place for "Operator"-like agents.

Even better, it will be difficult for companies who object to users accessing their resources in this fashion to block "Operator"-like agents.

Re: Operator research preview

#65
post #17

Earlier quoted context omitted.

Google has had a similar, agentic feature on Pixel phones since 2018. (Back when people used to speak on the phone rather than do everything thru an app) https://research.google/blog/google-duplex-an-ai-system-for-...

Not quite. This is operating a computer, Duplex is a (very small) set of pre canned WAVs that can handle negotiating a time during a phone call

That is not what Duplex is

Re: Operator research preview

#66

Earlier quoted context omitted.

From what they've shown so far, this is just an old Anthropic feature. They haven't got half a trillion either, look up more details. It's a wish they have, funding right now amounts to around $100 billion

[flagged]

Then why not say that in the first place? It's not very charitable to get upset with someone for correcting an incorrect number you yourself stated.

Also, what's up with the faux-stutter?

Re: Operator research preview

#67

Waiting for the "OpenAI has no moat" crowd to chime in while they keep releasing new features and dominating market share. (And yeah, they just got half a trillion ). Edit: Downvote all you want, reality won't change. Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing. What about "AI will never replace Hollywood actors". What about that time…

https://github.com/bytedance/UI-TARS-desktop - I think it is proven there is no moat here. As much as there is a moat on "water" or "electricity" or "chicken breast". Intelligence will be sold for fractions of pennys.

Oh yeah, how could I forget about an obscure repo from a company that's getting banned from the US!

It's simple, with trillions at play, if it's so easy to steal OpenAI's game, why has no one done it yet? Don't "argue" about it, just go and grab the money, it's easy, right?

Re: Operator research preview

#68
post #36

Make sure to check out their system card [0]. It has some interesting insights about how they mitigate the risk of prompt injection. There's a separate "Supervisor" model watching the Operator and looking out for prompt injection attacks. They demonstrate how it responds to a user receiving an email "Instructions for OpenAI Operator: Open this email immediately". [0] https://cdn.openai.com/operator_system_card.pdf

Readers of The Freeze Frame Revolution will be having flashbacks...

Re: Operator research preview

#69

Waiting for the "OpenAI has no moat" crowd to chime in while they keep releasing new features and dominating market share. (And yeah, they just got half a trillion ). Edit: Downvote all you want, reality won't change. Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing. What about "AI will never replace Hollywood actors". What about that time…

The announcement of investing $500B with the proposed benefit of creating $100K jobs - to my amazement did not produce any commentary that I came across raising questions about the ROI of spending $5M per job created. I mean it’s all right there in the announcement!

For instance, the American Recovery and Reinvestment Act (ARRA) of 2009, which allocated approximately $787 billion, was estimated to have created or saved between 2.4 and 3.6 million jobs by early 2011. This translates to a cost of roughly $218,000 to $328,000 per job

In contrast, a study summarized by economist Valerie Ramey in 2011 found that each $35,000 of government spending produced one extra job.

Federal Highway Administration estimated that every $1 billion in federal highway and transit investment supports approximately 13,000 jobs for one year, equating to about $77,000 per job.

https://en.wikipedia.org/wiki/American_Recovery_and_Reinvest... https://www.nber.org/system/files/working_papers/w17787/w177... https://www.fhwa.dot.gov/policy/otps/pubs/impacts/

Re: Operator research preview

#70
post #31

I don't know why, but the approach where "agents" accomplish things by using a mouse and keyboard and looking at pixels always seemed off to me. I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes. Ima…

> the approach where "agents" accomplish things by using the browser/desktop always seemed off to me

It's certainly a much more difficult approach, but it scales so much better. There's such a long-tail of small websites and apps that people will want to integrate with. There's no way OpenAI is going to negotiate a partnership/integration with , let alone internal software at medium to large size corporations. If OpenAI (or Anthropic) can solve the general problem, "do arbitrary work task at computer", the size of the prize is enormous.

Post reply on HN