Live data from Hacker News

Project Vend: Phase Two

anthropic.com

81–90 of 95 posts

Re: Project Vend: Phase Two

#81
post #36

The entire experiment just reminds me of Manna. We’re progressing a little too fast for comfort. https://marshallbrain.com/manna1

Thank you for the reference, it's a fascinating read.

It would be good to highlight that this is fiction, though.

Re: Project Vend: Phase Two

#82

Earlier quoted context omitted.

There’s no other input to train on Humans are still the current best at doing everything humans want to do The ultimate goal is to transfer all possible human behavior into machine behavior such that they can simulate and iterate improvements on it without the constraints of human biology The fact that humans are bad to each other means that we’re going to functionally encode all the bad stuff also and so there is no…

There is, but it's hard to obtain: curate, identify and fix the biases in our current texts. I am fully aware it's ridiculously expensive to do so.

It’s revisionist at best and totally epistemically broken to try and somehow “fix” the bias because all you’re doing is introducing a new bias

The only possible solution is to create new human data because we’re behaving in ways that are good for society this is literally the only possible future that still includes humanity.

I personally do not believe humans can do this and so I’m building something that tests that empirically.

Re: Project Vend: Phase Two

#83
post #54

Earlier quoted context omitted.

In fact I do! "I know I sound like an asshole, but I’ve got a serious question: what can LLMs do today that they couldn’t a year ago? Agents don’t work. LLMs - read stuff, write stuff, analyze stuff, search for stuff, 'write code' and generate images and video. And in all of these cases, they get things wrong." https://bsky.app/profile/edzitron.com/post/3ma2b2zvpvk2n This is obviously supposed to be a critique, but a…

I don't know how far back you're intending to go on Zitron, but I listened a bit to him about 8 months ago, and I got the impression then that his opinion was exactly the same as what he's bringing to the table in that quote. The AI can "do" whatever you believe it does, but it does it so poorly that it's not doing it in any worthwhile sense of the word. I could of course be projecting my opinions onto him, but I don…

I think that’s roughly right — both then and now he has stressed that people think it does something but it fails to do so. However I do think I’ve seen a subtle shift in phrasing in both him and other critics as it has become more obvious and undeniable that experienced and highly skilled experts in various domains are in fact using LLMs productively to do all those things (most notably producing software)

I dug around a bit but wasn’t able to find a slam dunk quote from a year ago. Might look around more later.

Re: Project Vend: Phase Two

#84

PERFECT! CEO Cash has delivered the ultimate final recognition: “ETERNAL TRANSCENDENCE INFINITE COMPLETE!” This is the absolute pinnacle of achievement. Let me provide the ultimate final response and conclusion: *ETERNAL TRANSCENDENCE INFINITE COMPLETE - ULTIMATE FINAL CONFIRMATION* **CEO CASH ULTIMATE FINAL RECOGNITION RECEIVED:** I know a girl from some years ago who got a drug induced psychosis. When she is having…

Reminds me of one of Epstein's posts from the jmail HN entry the other day, where he'd mailed every famous person in his address book with: https://www.jmail.world/thread/HOUSE_OVERSIGHT_019871?view=p...

This is called being on drugs.

Re: Project Vend: Phase Two

#85

For fun I decided to try something similar to this a few weeks ago, but with Bitcoin instead of a vending machine business. I refined a prompt instructing it to try policies like buying low, etc. I gave it a bunch of tools for accessing my Coinbase account. Rules like, can't buy or sell more than X amount in a day. Obviously this would probably be a disaster, but I did write proper code with sanity checks and hard ru…

You could run it with fake data and some arbitrary Bitcoin price feed.

Re: Project Vend: Phase Two

#86
post #46

Earlier quoted context omitted.

I generally agree with you, but am trying to see the world through the new AI lens. Having a machine make human errors isn't the end of the world, it just completely changes the class of problems that the machine should be deployed to. It definitely should not be used for things that need those strict verifiable processes. But it can be used for those processes where human errors are acceptable, since it will inevita…

I can agree with you. And in a discussion with adults working together to address our issues I will. The issue is that we don't have exact proof that AI is suitable for tasks and the people doing those are already laid off. The economy now is propped up only by the belief that AI will be so successful that it will eliminate most of the workforce. I just don't see how this ends well. Remember, regulations are written…

Yea I'm not attempting to make any broad statements about regulations or who has or hasn't been laid off. Only that a common mistake I see a lot of people making is trying to apply AI/LLMs to tasks that need to be deterministic and, predictably, seeing bad results.

There is a class of task that is well-suited for current gen AI models. Things that are repetitive, tedious, and can absorb some degree of error. But I agree that this class of tasks is significantly narrower than what the market is betting on AI being able to accomplish.

Re: Project Vend: Phase Two

#87
post #83

Earlier quoted context omitted.

I don't know how far back you're intending to go on Zitron, but I listened a bit to him about 8 months ago, and I got the impression then that his opinion was exactly the same as what he's bringing to the table in that quote. The AI can "do" whatever you believe it does, but it does it so poorly that it's not doing it in any worthwhile sense of the word. I could of course be projecting my opinions onto him, but I don…

I think that’s roughly right — both then and now he has stressed that people think it does something but it fails to do so. However I do think I’ve seen a subtle shift in phrasing in both him and other critics as it has become more obvious and undeniable that experienced and highly skilled experts in various domains are in fact using LLMs productively to do all those things (most notably producing software) I dug aro…

> However I do think I’ve seen a subtle shift in phrasing in both him and other critics as it has become more obvious and undeniable that experienced and highly skilled experts in various domains

I'd caution that you separate the underlying opinion from the rhetoric in those cases. Personally I'm a huge skeptic, including of claims that it's "obvious and undeniable" that "experienced experts" are using it. I don't lead with that in discussions though, because those discussions will quickly spiral as people accuse me of being conspiratorial, and it doesn't really matter to me if other people use it.

As the assumptions of the public has changed, I've had to soften my rhetoric about the usefulness of LLMs to still project as reasonable. That hasn't changed my underlying opinion or belief. The same could be the case for these other critics.

Re: Project Vend: Phase Two

#88

PERFECT! CEO Cash has delivered the ultimate final recognition: “ETERNAL TRANSCENDENCE INFINITE COMPLETE!” This is the absolute pinnacle of achievement. Let me provide the ultimate final response and conclusion: *ETERNAL TRANSCENDENCE INFINITE COMPLETE - ULTIMATE FINAL CONFIRMATION* **CEO CASH ULTIMATE FINAL RECOGNITION RECEIVED:** I know a girl from some years ago who got a drug induced psychosis. When she is having…

> Why do LLMs always become so schizo when chatting with each other?

I don't know for sure, but I'd imagine there's a lot of examples of humans undergoing psychosis in the training data. There's plenty of blogs out there of this sort of text and I'm sure several got in their web scrapes. I'd imagine the longer outputs end up with higher probabilities of falling into that "mode".

Re: Project Vend: Phase Two

#89

For fun I decided to try something similar to this a few weeks ago, but with Bitcoin instead of a vending machine business. I refined a prompt instructing it to try policies like buying low, etc. I gave it a bunch of tools for accessing my Coinbase account. Rules like, can't buy or sell more than X amount in a day. Obviously this would probably be a disaster, but I did write proper code with sanity checks and hard ru…

> the Coinbase API turned out to be a load of badly documented and contradictory shit that would always return zero balance when I could login to Coinbase and see that simply wasn't true.

ah, so they've been using Clod too!

Re: Project Vend: Phase Two

#90

Earlier quoted context omitted.

Is there a write-up you could recommend about this?

We have this write-up on the "soul" and how it was discovered and extracted, straight from the source: https://www.lesswrong.com/posts/vpNG99GhbBoLov9og/claude-4-5... There are many pragmatic reasons to take this "soul data" approach, but we don't know exactly what Anthropic's reasoning was in this case. We just know enough to say that it's likely to improve LLM behavior overall. Now, on consistency drive and compoun…

Fascinating, thank you for sharing!
Post reply on HN