Live data from Hacker News

Project Vend: Phase Two

anthropic.com

51–60 of 95 posts

Re: Project Vend: Phase Two

#51
post #3

Earlier quoted context omitted.

We will poor billions into this until you are begging for us to run your business!

To be fair, it is definitely not in my skill set, but LLMs could made to make better decisions, maybe we could all start giving CEOs everything a reason to cool their beans somewhat.

A lot of work in project controls and management are simple enough that any system that can handle data that isn’t reliably structured could do it. Read project team updates each week. Are we on time and on budget? If yes, commend the team and write a glowing report of the AI’s wise and dynamic leadership to operations, if not, encourage the team and recommend operations outsource the employees.

Re: Project Vend: Phase Two

#53

To me the key point was: > One way of looking at this is that we rediscovered that bureaucracy matters. Although some might chafe against procedures and checklists, they exist for a reason: providing a kind of institutional memory that helps employees avoid common screwups at work. That's why we want machines in our systems - to eliminate human errors. That's why we implement strict verifiable processes - to minimize…

I don’t think they really want to fix human errors with LLMs. Rather, they want a “human” who works 24x7 for dirt cheap

Re: Project Vend: Phase Two

#54
post #25

This is a great read. I just want to point out what great marketing this and the WSJ story are. People reading it think they’re sticking it to Anthropic by noticing that Claude is not that good at running a business, meanwhile the unstated premise is reinforced: of course Claude is good at many other things. I have seen a shift in the past few months among even the most ardent critics of LLMs like Ed Zitron: they’ve…

Zitron has never said anything like that. Do you have a quote?

In fact I do!

"I know I sound like an asshole, but I’ve got a serious question: what can LLMs do today that they couldn’t a year ago? Agents don’t work. LLMs - read stuff, write stuff, analyze stuff, search for stuff, 'write code' and generate images and video. And in all of these cases, they get things wrong."

https://bsky.app/profile/edzitron.com/post/3ma2b2zvpvk2n

This is obviously supposed to be a critique, but a year ago he would never have admitted LLMs can do any of these things, even with errors. This seems strange but it's typical of Zitron's writing, which is often incoherent in service of sounding as negative as possible. A couple of other examples I've written about are his claims about the "cost of inference" going up and about Anthropic allegedly screwing over Cursor by raising prices on them:

https://crespo.business/posts/cost-of-inference/

https://news.ycombinator.com/item?id=45645714

Re: Project Vend: Phase Two

#55

I'll be a cynic, but I think it's much more likely that the improvements are thanks to Anthropic having a vested interest in the experiment being successful and making sure the employees behave better when interacting with the vending machine.

I wonder if it's the opposite actually. When there is a human running a convenience store type of thing, people don't generally spend time trying to convince them of obviously absurd things, particularly if they work for the same company as you. Nobody wants to risk the employee refusing to sell anything to you because you're a time-wasting jerk or maybe their manager telling them to stop wasting time messing with their co-worker.

Re: Project Vend: Phase Two

#56

Earlier quoted context omitted.

These kind of agents really do see the world through a straw. If you hand one a document it doesn't have any context clues or external methods of determining its veracity. Unless a board-meeting transcript is so self-evidently ridiculous that it can't be true, how is it supposed to know its not real?

I don't think it's that different to what I observe in humans I work with. Things that happen regularly (and I have no reason will change in the future): 1) Making the same bad decisions multiple times, and having no recollection of it happening (or at least pretending to have none) and without any attempt to implement measures to prevent it from happening in the future 2) Trying to please people (I read it as: tryin…

I don't think it's that different to what I observe in humans I work with.

If the "AI" isn't better at its job than a human, then what's the point?

Re: Project Vend: Phase Two

#57

other than these tests I actually rarely see vending machines. are they really representative or popular still in usa?

other than these tests I actually rarely see vending machines. are they really representative or popular still in usa?

I guess you've never been to Asia, either.

It's a big world.

Re: Project Vend: Phase Two

#59

The cynicism is wild - there is a computer running a store largely autonomously. I can’t imagine being interested in computers and NOT finding this wildly amazing

Its very much NOT running a store. They took an employee fridge and added a LARP to it.

That they are framing this as a legitimate business is either misunderstanding their current position in the economy, or deliberate misdirection. We're not playing around with role playing chatbots anymore. This shit was supposed to be displacing actual humans.

Re: Project Vend: Phase Two

#60

Earlier quoted context omitted.

I don't think it's that different to what I observe in humans I work with. Things that happen regularly (and I have no reason will change in the future): 1) Making the same bad decisions multiple times, and having no recollection of it happening (or at least pretending to have none) and without any attempt to implement measures to prevent it from happening in the future 2) Trying to please people (I read it as: tryin…

I don't think it's that different to what I observe in humans I work with. If the "AI" isn't better at its job than a human, then what's the point?

Idk, seems like a different topic, no?

Off the top of my head, things that could be considered "the point":

- It's much cheaper

- It's more replicable

- It can be scaled more readily

But again, not what I was arguing for or against; my comment mostly pertained to "world through a straw"

Post reply on HN