We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
211–220 of 258 posts
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#212The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…
This is HN for Christ's sake. Stop treating deterministic algorithms like they are humans.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#213The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…
But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM.
I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM.
In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision."
Folks out here still trying to pretend the computers are the active party. They are not.
Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#214Earlier quoted context omitted.
That's an incredibly deep misunderstanding. Almost as bad as saying that human is the same as a tree because we're both made of carbohydrates and proteins.
The comparison I provided is between how an LLM functions and one part of how a brain functions. It's not an equivalence, I did not say they are "the same". You made the claim that these systems "aren't remotely comparable", and when faced with a clear comparison, you claim "deep misunderstanding".. Have you any arguments to make, or is this going to devolve into more statements that both mischaracterize and muddy th…
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#215The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…
This is HN for Christ's sake. Stop treating deterministic algorithms like they are humans.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#216Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#217Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#218Earlier quoted context omitted.
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#219Earlier quoted context omitted.
100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives
I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?" The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide." Even a low quality local thinking model that has been tuned to be unhin…