Live data from Hacker News

Operator research preview

openai.com

151–160 of 448 posts

Re: Operator research preview

#151

Earlier quoted context omitted.

As the storyline unfolds "AI" seems to be code for "machine learning based censorship". Soon we will have home appliances and vehicles telling you about how aligned you are, and whether you need to improve your alignment score before you can open your fridge. It is only a matter of time before this will apply to your financial transactions as well.

I can sympathize with vague notions of AI dystopia, but this might be stretching the concept a bit too far. This kind of service is extremely abusable ("Operator, go to Wikipedia and start mass-vandalizing articles" or "Go to this website and try these people's email addresses with random passwords until it locks their accounts") and building some alignment goals into it doesn't seem like a terribly draconian idea. A…

You can also write a python script to achieve the same goals.

Except it's not python's responsibility to interpret the intent of your script, just as it's not your phone's responsibility to interpret the contents of your conversation.

So our tools are not our morality police. We have a legal system that can operate within the bounds of law and due process. I am well aware of the already applied levels of machine learning policing, I am just not very excited that society has decided that "this is the way now", and also doesn't seem to be bothered by the environmental costs of building and running all these GPUs (which does seem to be the case when they are used for censorship resistant transactions), or the ethical concerns about a non-profit becoming a for-profit etc.

Re: Operator research preview

#152

Earlier quoted context omitted.

As the storyline unfolds "AI" seems to be code for "machine learning based censorship". Soon we will have home appliances and vehicles telling you about how aligned you are, and whether you need to improve your alignment score before you can open your fridge. It is only a matter of time before this will apply to your financial transactions as well.

I can sympathize with vague notions of AI dystopia, but this might be stretching the concept a bit too far. This kind of service is extremely abusable ("Operator, go to Wikipedia and start mass-vandalizing articles" or "Go to this website and try these people's email addresses with random passwords until it locks their accounts") and building some alignment goals into it doesn't seem like a terribly draconian idea. A…

I don't think webmasters will be sitting down and hoping that this will not be abusable. Unlikely these kinds of agents would be allowed at all for producing content of any kind automatically (e.g. not via their APIs), or ai-slop will just overwhelm the internet exponentially.

The same neural networks are ready for detecting certain fingerprints and denying them entrance

Re: Operator research preview

#153

Waiting for the "OpenAI has no moat" crowd to chime in while they keep releasing new features and dominating market share. (And yeah, they just got half a trillion ). Edit: Downvote all you want, reality won't change. Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing. What about "AI will never replace Hollywood actors". What about that time…

Why cheer for a private company that does not care about anyone and wants to replace the internet with their crappy interface?

Re: Operator research preview

#154

Earlier quoted context omitted.

Big jumps in benchmarks from Claude's Computer Use though. 87% vs 56% on Webvoyager 58.1% vs 36.2% on WebArena 38.1% vs 22% on OsWorld These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)

OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-use agents (which can use the DOM, browser-specific APIs, and so on.) The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with…

Yeah, and Browser Use already has 89% on WebVoyager https://browser-use.com/posts/sota-technical-report

Re: Operator research preview

#155
What is fascinating about this announcement is if you look into future after considerable improvements in product and the model, we will be just chatting with ChatGPT to book dinner tables, flights, buy groceries and do all sort of mundane and hugely boring things we do on the web, just by talking to the agents. I'd definitely love that.

Re: Operator research preview

#157

From the slide deck on the livestream: "[Operator safety risks and mitigations] Harmful tasks: User is misaligned" Looking forward to seeing some more of the examples for when openai considers their users as "misaligned", whatever that actually even means anymore.

OAI has decided to stop aligning models and focus on aligning the users instead.

"Society is fixed, biology is mutable", but taken to the extreme?

Re: Operator research preview

#158
post #60

Overall, Operator seems the same as Claude's Computer Use demo from a few months ago, including architecture requiring user to launch a VM, and a tendency to be incorrect: https://news.ycombinator.com/item?id=41914989 Notably, Claude's Computer Use implementation made few waves in the AI Agent industry since that announcement despite the hype.

I thought Claude Computer Use is through API, and I remember hearing about high number of queries and charges. This looks like its in browser through the standard $20 Pro fee, which is huge. (EDIT: $200 a month plan so less of a slam dunk but still might be worth it) Is there any open source or cheap ways to automate things on your computer? For instance I was thinking about a workflow like: 1. Use web to search for…

If you are worried about costs you can use Browser Use with deepseek which becomes super cheap! https://github.com/browser-use/browser-use
Post reply on HN