Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

131–140 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#133
post #66

Earlier quoted context omitted.

i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.

Fail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well. “Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.

maybe it is because I am biased but I have almost no expectation for AI to act "ethically"

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#134

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

[deleted]

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#135
post #82
post #63

Earlier quoted context omitted.

The promise of AI: unlimited power. I mean spam. Unlimited spam.

Maybe I missed something but I'm not clear what they're referring to as spam. I guess the fact that the agent emailed all users with discounts and dropped the price a few times? I don't think that's usually what people call spam. (For example if it had emailed everyone once would we call that spam? No. So it's about frequency of price drops?)

They did include a screenshot which looks like at least 6 emails being sent in the 24 hour time window. I would certainly consider that spamming from some diary app on my phone.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#136

Earlier quoted context omitted.

You can literally read their thoughts if you run an open model, they look like pretty human thoughts to me, albeit a neurotic human.

These aren't thoughts how humans literally think them. I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.

> aren't remotely comparable to the way humans think and act

Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#137

It seems the agent was stymied by being bot blocked so often. I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.

That is hilarious, depressing, and would likely work.

Oh god.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#138

Earlier quoted context omitted.

> Results that arrive after the deadline do not exist Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)

Sounds like every startup I ever worked for. What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever

> What’s the line?

Evidence, even when downplayed or ignored, is still evidence.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#139

Earlier quoted context omitted.

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"

An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does? If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're…

They're things we are intentionally engineering in our own image, based on massive statistical analysis of our own actions and behavior. So what's there to not understand? If this wasn't the case, that would be much weirder.

They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.

Post reply on HN