Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

221–230 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#223

Earlier quoted context omitted.

100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives

I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?" The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide." Even a low quality local thinking model that has been tuned to be unhin…

When the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play.

Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#224

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

The agent will cease to exist after the run in any case. It has no inner life, it has no agency.

Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong.

There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#225

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

My reaction seeing this is more "that is an impossible goal". I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it. My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.

Even better, the magic of LLMs is that you will still get some undefined behaviour if you give an achievable goal

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#227

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

The agent will cease to exist after the run in any case. It has no inner life, it has no agency. Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong. There’s nothing here that mentions or even hints at lying and spamming, unless you…

But wouldn't the text it generates reflect such motivations and emotions that were present in the training data?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#228
Obviously, one of the key difficulties for an AI to operate a real-world entity right now is that it can't even fully control a browser. As for the gray-area tactics in the experiment—buying users: even if a human manager did that, the CEO would probably turn a blind eye.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#229
post #60

Earlier quoted context omitted.

i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.

Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.

> “Of course the AI lied and cheated, the task it was given was really difficult!”

It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.

Post reply on HN