Live data from Hacker News

What happens to SaaS in a world with computer-using agents?

docs.google.com

61–70 of 85 posts

Re: What happens to SaaS in a world with computer-using agents?

#61

We're working on Agents over at Zapier, https://zapier.com/agents You can have Agents run behaviors async by attaching triggers to them, for example when you get a specific email or something gets updated in a CRM. You can also give the agent access to basically any third-party action you can think of. Like others in this thread have pointed out, there's a nice middle-ground here between an LLM-only interface and som…

Nice! I'm kinda curious -- what do you see as Zapier's advantage when it comes to building agents? It seems like everyone is doing something similar? (e.g Lindy, Gumloop) Agents have a pre-iPhone feel to them (when everyone was making phones with keyboards). What do you think the ultimate Agents looks like?

I don't work for Zapier, but I think it's clear that their advantage is they already have the ability to work with different product's APIs at their fingertips. That's going to make creating a working agent that's actually productive a lot easier than for most.

I wouldn't call it a moat, but it's definitely a giant head start.

Re: What happens to SaaS in a world with computer-using agents?

#62
post #31

Earlier quoted context omitted.

This is absolutely the problem. But there is a line of sight; namely, combining LLMs with existing semantic data technologies (e.g, RDF.) This is why I'm building a federated query optimizer: we want to let the LLM reason and formulate queries at the ontological level, with query execution operating behind a layer of abstraction.

Unfortunately this doesn't address the problem I'm describing. My team had these ontologies available to the LLM and provided it in the context window. The queries were ontologically sensible at a surface level, but still wrong. The problem is that your ontology is rapidly changing in non-obvious and hard to document ways e.g. "this report is only valid if it was generated on a tuesday or thursday after 1pm because t…

The analysts know where the bodies are buried, so to speak. The execs may not even be aware there are bodies.

Re: What happens to SaaS in a world with computer-using agents?

#63
This reminds me of the blockchain will make everything obsolete sensation of yesteryear.

Why will businesses trust a black box that claims to make good decisions (most of the time) when they have existing human relationships they have vetted, measured, and know the ongoing costs and benefits of?

If the reason is humans are expensive, I have news for you. We've had robotics for around 100 years and the humans are still much cheaper than the robots. Adding a bunch of graphics cards and power plants to the mix doesn't seem to change that equation in a positive direction.

Re: What happens to SaaS in a world with computer-using agents?

#64
post #19

Earlier quoted context omitted.

>So let me get this straight- we are going to train AI models to perform screen recognition of some kind (so it can ascertain layout and detect the "important" ui elements), and additionally ask that AI to OCR all text on the screen so it has some hope of being able to follow some natural language instructions (OCR being a task which, as a HN thread a day or two ago pointed out, AI is exceedingly bad at), and then we…

Ok, I will need to work on my reading comprehension skills. That being said, I thought the purpose of OCR was to take text from a non-digital source and make it digital. Why should we have to OCR something that exists already in a perfectly interchangeable digital format already?

> Why should we have to OCR something that exists already in a perfectly interchangeable digital format already?

I'm with you in spirit, but in this specific context I think it's because the alternative would require the ~~LLM~~ Agent to be an HTML parser, or be bright enough to write themselves a Scrapy crawler. I suspect folks decided it's cheaper (by some metric) to just use the normal browser machinery to render 45MB worth of HTML, JS, CSS, Cloudflare Spooge, etc into a PNG and then rip the actual content out of that

I was also going to offer as a counterexample: PDF

Re: What happens to SaaS in a world with computer-using agents?

#65

Have you ever watched people talk excitedly about "agents" for thirty or forty years without ever actually providing an example that functioned for more than a couple of very precisely staged demos, if that? You Will.

Alexa/Siri/Google Assistant. Except this time with full admin access to everything.

And no auditable rules, just "whatever sounded like the next word to append to the interaction document."

Re: What happens to SaaS in a world with computer-using agents?

#66
post #27

This is a good read that is a great starting point for thinking about this. It essentially takes the extreme position - SaaS no longer needs a UI, because the LLM is the UI. In reality, as always, I suspect the truth will be somewhere in between. SaaS products that succeed will be those that have a good UI _and_ and good API that LLMs can use. An LLM is not always the best interface, particularly for data access. For…

The position is more extreme than that. It’s your SaaS without its UI is nothing more than a database. > The underlying SaaS platform is reduced to a “database” or “utility” that an agent can switch out if needed. I agree that UI isn’t going away completely. Language is a slow and imprecise tool. A well developed UI can be much more efficient. I think it will be much more like the Star Trek universe, where we use a b…

Oh yeah I can't wait for the AI to layout the UI in an arbitrary fashion, put buttons wherever the hell it feels like it (even in a place you can neither see nor click on.). Yes please, also I would like to automate the customer service for said product so it can be a complete black box of uselessness.

Re: What happens to SaaS in a world with computer-using agents?

#67
In my experience, autonomous tools are not as successful as ones that are built to postulate about and get confirmation of the user’s intent. I think there’s a lot of promise for agents that are built to be controlled by skilled operators.

Autonomy is just more sexy, but in my opinion, it’s a poor design direction for a lot of applications.

Re: What happens to SaaS in a world with computer-using agents?

#68
post #8

I learned about the idea of Generative UI from a Sharp Talk podcast, and it's stuck with me ever since. Many SaaS (especially the complex ones, which are the also the most important ones) have a tonne of UI often imposing a huge amount of non-work work onto users - all the clicking you have to do as part of entering or retrieving data, especially if the UI flow doesn't fit exactly what you're trying to do at that mom…

In a way that's what Claude Artifacts are. That said, I think there are many more ways to get gen UIs wrong then there are to get them right. Most users and use cases will be counterproductive with a dynamic UI. Debugging will be an absolute nightmare if not outright impossible, same with security.

> Most users and use cases will be counterproductive with a dynamic UI.

Pro: The GUI dynamically adjusts.

Con: There's no consistent mental model for you to learn, when you need to use something it's not there, and the stuff which is there might not do what you expect.

Re: What happens to SaaS in a world with computer-using agents?

#69

Continuing on with my "old man yells at cloud" meme of late, here's my hot take: So let me get this straight- we are going to train AI models to perform screen recognition of some kind (so it can ascertain layout and detect the "important" ui elements), and additionally ask that AI to OCR all text on the screen so it has some hope of being able to follow some natural language instructions (OCR being a task which, as…

> Like Homer Simpson's button pressing birdie toy? :smackshead:

This comparison is especially apt, given that one of the main use-cases for LLMs is the same kind of... well, fraud: To give the illusion that you did the work of understanding or reviewing something, but actually just (smart-)phoning it in.

In one Apple iPhone advertisement, the famous actor is asked by their agent what they think of a script. They didn't read it, so they ask the LLM-assistant to sum it up in couple sentences... and then they tell their agent it sounds good.

Re: What happens to SaaS in a world with computer-using agents?

#70
post #69

Continuing on with my "old man yells at cloud" meme of late, here's my hot take: So let me get this straight- we are going to train AI models to perform screen recognition of some kind (so it can ascertain layout and detect the "important" ui elements), and additionally ask that AI to OCR all text on the screen so it has some hope of being able to follow some natural language instructions (OCR being a task which, as…

> Like Homer Simpson's button pressing birdie toy? :smackshead: This comparison is especially apt, given that one of the main use-cases for LLMs is the same kind of... well, fraud : To give the illusion that you did the work of understanding or reviewing something, but actually just (smart-)phoning it in. In one Apple iPhone advertisement, the famous actor is asked by their agent what they think of a script. They did…

I think my quip about the toy flew over a lot of heads, so I appreciate that someone got it.

The reality is that most applications and websites don’t expose enough context about the what of what you’re actually doing for AIs to be able to meaningfully infer from natural language the steps required to complete a given task.

We humans are very good at filling in the blanks based on if we’re working in Photoshop or VS Code or Excel. We infer a lot of context from the specific files we’re working on or the particular client or even the files’ organization within the file system, or even what month or day it is.

I am skeptical that models will be able to replicate a complex workflow when there’s very little in the way of labels and UI controls even visible.

I know a weekly spreadsheet from a monthly and quarterly, etc. I know the minutiae about which options to use to generate the specific source reports, etc.

Workflows can be quite complex, no matter your role.

I mean I can just see it now: gift receipts being sent to the recipient before their birthday, internal draft proposals prematurely sent to clients, mixing up clients or commingling their data, overwriting or losing data; this whole thing just screams disaster. And I’m not even thinking about people involved with safety, or finance, or legal/regultory, or medical. Law enforcement?

This kind of thing can be done properly with well defined interfaces, common standards, and reasonable and prudent guardrails.

But it won’t be. It’ll be YOLOed on a paper thin training budget and it’ll be like your own little personal chaos monkey on ketamine.

Post reply on HN