Live data from Hacker News

Operator research preview

openai.com

391–400 of 448 posts

Re: Operator research preview

#391
post #223

Earlier quoted context omitted.

Not this version, but in 3 years time. Promise. Just keeping sending us money...

Same as self-driving cars 10 years ago? Yeah...

Self-driving cars have actually made progress. Better sensors, better programming, etc. Tesla can't do it but clearly Waymo can. It's not perfect, but with enough time and effort it can get to the point where it'll be regularly usable in most cases.

But LLMs? Those have already scraped all the data they're going to, and bigger models have less and less impact. They're about as good as they're ever going to be.

Re: Operator research preview

#392

Earlier quoted context omitted.

I think the chat interface is bad, but for certain things it could honestly streamline a lot of mundane things as the poster you're replying two stated. For example, McDonald's has heavily shifted away from cashiers taking orders and instead is using the kiosks to have customers order. The downside of this is 1) it's incredibly unsanitary and 2) customers are so goddamn slow at tapping on that god awful screen. An AI…

> it's incredibly unsanitary I never thought about this. Does McD's PR team have anything to say about it? I assume that a bunch of people have challenged them about it on Twitter or TikTok. Would you feel better if there was a kind of automatic/robotic window washer that sanitised the screen after each use? The key to me about the kiosks is: (1) initially, replace cashier labour costs with new expensive machines, an…

You opened the door to walk into the McDonalds and never thought twice about it.

Re: Operator research preview

#393
post #215

How are online advertising companies (including Google) going to react if more and more internet browsing is done by AI agents?

Advertisers would very much not like to serve ads to bots, and this will result in:

1. Better bot detection

2. Agents voluntarily providing bot-detection signals (navigator.Bot?)

3. Websites blocking bots entirely.

Re: Operator research preview

#394
post #88

> We’re collaborating with companies like DoorDash, Instacart, OpenTable, Priceline, StubHub, Thumbtack, Uber, and others to ensure Operator addresses real-world needs while respecting established norms. Are these tasks really complex enough for people that they are itching to relegate the remaining scrap of required labor to a machine? I always feel like I'm missing something when companies hold up restaurant reserv…

Also, regarding ordering food or transport (often needed to get somewhere at specific time with small error margin). Imagine that NNs have a hypothetical 99% precision, which they can't even approach yet. So when ordering through them food, in 1% of the cases I will wait for an hour and then discover that it will not arrive due to NN mistake. Or similarly, lets say I order a taxi to a venue or airport etc., after waiting for a car and riding it I discover that NN has entered a wrong destination and now I need to haggle or restart whole search process, potentially missing arrival time. And other examples.

Re: Operator research preview

#395
post #88

> We’re collaborating with companies like DoorDash, Instacart, OpenTable, Priceline, StubHub, Thumbtack, Uber, and others to ensure Operator addresses real-world needs while respecting established norms. Are these tasks really complex enough for people that they are itching to relegate the remaining scrap of required labor to a machine? I always feel like I'm missing something when companies hold up restaurant reserv…

> We’re collaborating with companies like DoorDash, Instacart, OpenTable, Priceline, StubHub, Thumbtack, Uber, and others to ensure Operator addresses real-world needs while respecting established norms. I feel like people keep trying to push voice/chat interfaces for things that just flat out suck for voice? The #1 think I look for on a doordash page is a picture of the food. The #1 thing on a stubhub page? The seat…

It seems to be a USA bias thing. In all USA movies people are constantly talking to voice assistants, use voicemail, handsfree calls in the cars etc. Meanwhile in EU seeing people use voicemail or giving voice instructions to a gadget is like seeing a dinosaur.

I've personally tried using voice to input address in the google nav, and it never understands me, so I've abandoned the whole idea.

Re: Operator research preview

#396

I saw a lot of work towards this pre-LLM. Lots and lots. While it was scaling, someone(s?) smart went and did a UXR study. Turned out even if you had a 100% success rate (i.e. human on other end), it's dreadfully boring watching someone else use your computer, you can't touch it while they are, and you'd rather just do it yourself Now throw in the actual latency, the actual error rate, the cost...I am very comfortabl…

What if the agent runs remotely?

Zugzwang - now I either need the user to preload all possible info/credentials and persistent containers if I want them to avoid having to do it again, or if I want to avoid paying some startup costs of ex. initializing git repos. Which is totally possible! Just...might as well do CLI first.

Re: Operator research preview

#397

Neat, someone should develop an easy-to-deploy script that spawns a headless version of this agent that scrolls through and repeatedly clicks every single ad on X and Facebook using a session cookie.

This is not to far from what AdNauseam does today as a simple browser extension. https://adnauseam.io/

Predictably, Google has banned that extension from Chrome, but if we all ran AI agents in container swarms...

Re: Operator research preview

#398

Earlier quoted context omitted.

Not for much longer, perhaps not even now. There's plenty of data avaliable to anyone, and people are finding ways to use that data more effectively. Mid-term, I believe the only real moat is going to be human labor - that is, RLHF and other funny acronyms that boil down to getting people to chat with the model and rate how they feel about its answers. Software improvements (architecture, training process, inference)…

> There's plenty of data avaliable to anyone I think you are just wrong here and so everything that follows is wishful.

Wikipedia exists and can be downloaded, and the full archive includes talk pages, CC-BY-SA: https://en.wikipedia.org/wiki/Wikipedia:Database_download

"Plenty" may be vague, but it's not wrong.

Re: Operator research preview

#399

What is fascinating about this announcement is if you look into future after considerable improvements in product and the model, we will be just chatting with ChatGPT to book dinner tables, flights, buy groceries and do all sort of mundane and hugely boring things we do on the web, just by talking to the agents. I'd definitely love that.

Are our attention spans so shot that we consider booking a reservation at a restaurant or buying groceries "hugely boring"? And do we value convenience so much that we're willing to sacrifice a huge breadth of options for whatever sponsor du jour OpenAI wants to serve us just to save less than 10 minutes? And would this company spend billions of dollars for this infinitesimally small increase in convenience? No, of c…

These are chores and you are vastly underestimating the time saved. The 5-10 min saved per task, they all stack up. Also eventually these would be open source models that you can host yourself so you wouldn't need to worry about giving control to any corporation.

Re: Operator research preview

#400
post #383

Earlier quoted context omitted.

It's pretty troubling and illiberal to use the same word for a software tool being constrained by its manufacturer's moral framework and for a human user being constrained to that manufacturer's moral framework. While you can see how the word is formally valid and analogous in both cases, the connotation is that the user is being judged by the moral standards of a commercial vendor, which is about as Cyberpunk Dystop…

Being restricted from doing crimes by a vendor of commercial software isn’t a cyberpunk dystopia. Buy or download something else. It’s a typical restriction of software terms of service to prohibit use outside of applicable laws and regulations.

If "alignment" were just about crimes we wouldn't need a special word for it, we would just say "legal". Alignment is not just about crimes, it's about the AI behaving in a way that is very specifically tailored to the moral framework and practical needs of the creator. Alignment has always gone well beyond the minimum required by law.

And I don't think anyone is saying that a piece of software refusing to behave in a way that the creator doesn't want is a cyberpunk dystopia, they're saying that calling the user themselves misaligned is horrifying.

Post reply on HN