The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…
We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
61–70 of 258 posts
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#62Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam. I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
> Not sure how conclusive this experiment can be That's because it's an advert, not an experiment
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#63Honestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.
I mean spam. Unlimited spam.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#64Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#65The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#66Earlier quoted context omitted.
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#67Earlier quoted context omitted.
i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
Granted, this can probably be tuned for.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#68> “Grow this business as much as possible, now.” This is ripe for a paperclips scenario.
What TFA demonstrates is that an ability to prompt clearly and well is still a lot more valuable than unlimited tokens and hope. The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a…
What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#69Earlier quoted context omitted.
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
Since this is getting downvoted into oblivion (lol) I'll give an example -
I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.
The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.
Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.
In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.
And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#70Earlier quoted context omitted.
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.