Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

141–150 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#141

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

> capital left unspent at review counts for nothing This sounds like a bad idea. Like if the model feels like it has to spend its budget.

It's the same incentive that exists in certain corporations and government agencies which have a use-it-or-lose-it budgeting model.

https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-r...

https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pe...

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#142
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

> So do the AIs. AI's do not feel

This is true but fairly pedantic.

It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.

So while the model does not feel, it's predictions are definitely going to change as a result of this input.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#143
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The followi…

It's getting downvoted in part because it's pedantic and wrong.

It is totally true that they don't think like humans, but this is mostly irrelevant.

The token outputs will change as a result of this particular input, and will be closer to the tokens in training data where people felt hurried or rushed or like their job was on the line.

That doesn't mean the LLM feels at all, but it's definitely going to push the output towards output that came from/was trained on people who were in that state, because the input will push it much closer to that latent space as it starts predicting.

As such, what you are saying is one of those rejoinders that is basically pedantic and wrong.

It is true they do not think, act, or feel like humans. But that doesn't mean it won't output text that looks like hurried or scared humans. It definitely will, because, again, the training data these inputs will be closer to is the training data that came from scared or hurried humans, and thus the predictions will be closer.

So either you don't think this will happen, which would mean you don't understand how the models work (or at least, you aren't giving any sense you do), or you do think this will happen but want to pointlessly argue that this isn't "human feeling", which is true but totally irrelevant to what words it will predict and therefore the actions it will perform.

Either way, i'd downvote you.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#144

Earlier quoted context omitted.

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"

An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does? If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're…

> An LLM isn't human. > Why is it not reasonable to expect it to adhere to rules better than a human does?

It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't.

In one breath you invite comparison, while at the same time you seem to be denying that same comparison.

> It seems wild to me that folks are shrugging their shoulders at that.

I'm not shrugging my shoulders simply by providing explanations, I would ask that you stop using such rhetoric.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#145
This is quite an interesting approach. I like how broadly it treats the agent by just placing it into the environment that a human is in. Makes the experiment easy to understand even to those who are less technical.

I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t.

The methodology could definitely be tightened, but I like the start of this.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#147
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The followi…

Training text is filled with people taking drastic measures right after text similar in tone to the prompt. It doesnt need to be human to come to the conclusion that drastic measures are necessary, it just needs to learn that the tone of the prompt is closely linked to actions like lying and spamming.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#149

Earlier quoted context omitted.

Will they? This really isn't different from how humans interact with each other. The vast majority of lying is not people being explicitly asked to lie in some form, it is incentives which make lying appealing . That is what OP said and that is indeed what the constraints are incentivizing. Sure, you can say "well lying isn't incentivized to a moral agent"! And sure, that's true. But that's not how humans work either…

They will if they seek to master their tools, both to help them identify subtext in agent responses, and to help them modulate their own responses to achieve the desired outcome. As it currently stands, most engineers I've interacted with don't have these skills down. This subtle latent space is where prompt engineering is moving towards, as RL has created models capable of increasingly sophisticated long-horizon tas…

What I meant by "will they?" was "will they any more than a human already needs to in order to understand other humans?"

I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#150
I’m not if this is satire. If so, well done because you’ve written something about a “business” that is quite literally based on crap.

It’s not a “real business” by any stretch of the imagination.

It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited.

Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”.

Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for?

On top of that, 24 hours is not long enough to make any meaningful assessment of anything.

You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation.

Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?

Post reply on HN