Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

171–180 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#171
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

AIs feel? Maybe language structure in trading documents that ultimately led to fraud. If the latter is the case maybe AIs should not be trained on “negative outcomes.” I do not think AIs have emotions or are pressured by language either written or physical, just tokens.

Of course it is just tokens, but the result is the same.

If, in the amount of data they ingested, there was a clear pattern of responding in an hasty and carefree way to frenetic questions, LLMs will try more hasty and carefree solutions to a frenetic prompt.

You can decide whether you can say that they "feel" the urgency or not, but the outcome is very much the same

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#172

Earlier quoted context omitted.

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"

An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does? If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're…

Broadly speaking, I agree with your frustration, but I think this specific case is different. LLMs respond strongly to tone in wording, because they are trained on wording, and wording often has flexible meaning depending on context.

It's not a stretch to imagine that the training would cause it to respond this way. It would, in fact, be a greater stretch to argue that an LLM has a universal model in which it understands the concept of lying and truth, and can be primed to only use one or the other unless explicitly instructed otherwise.

After all, LLMs lie every time they tell you to run a command with bad arguments, or spit out some code with syntax errors.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#173

Earlier quoted context omitted.

They will if they seek to master their tools, both to help them identify subtext in agent responses, and to help them modulate their own responses to achieve the desired outcome. As it currently stands, most engineers I've interacted with don't have these skills down. This subtle latent space is where prompt engineering is moving towards, as RL has created models capable of increasingly sophisticated long-horizon tas…

What I meant by "will they?" was "will they any more than a human already needs to in order to understand other humans?" I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).

It's probably true that many programmers in the future will get away with a similar lack of fundamental knowledge that today's programmers get away with. To some degree, we all have blind spots, but I think if agentic processes are here to stay, as long as humans remain in the loop at all it would serve us to master a semantic capability closer to that of the models we work with, lest we lose control either in taste or in a manner more serious. The most effective engineers will understand that.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#174

Earlier quoted context omitted.

I feel like new graduates will need to start taking linguistics, psychology and public speaking classes in order to understand why and how subtext matters, and how to control it. Then again, we might find newer generations just develop an intuition in the same way that I witness some toddlers interface with touchscreens better than their parents.

You're expecting the vast majority of users for the deskilling machine to somehow want to learn a complicated subject then practice to get better at the subject by talking intricate classes and dedicating substantial amount of hours to learn how to better communicate with the deskilling machine? Hopefully these aren't the same graduates that just cheated their way through university, only the responsible users of LLM…

I don't think we can use the climate of today as indication of what comes tomorrow. Too much is in flux, we are experiencing growing pains. Few predicted what would happen to the world wide web in the early 90s, both the good and bad.

Plenty people today allow the internet to be a detrimental factor in their lives and don't have good habits built around it. The same will be true of AI.

However, we don't know what kind of engineering jobs will be left in one decade, much less two or three. Mastery may become generally important, or at least still be the difference between an adequately-compensated engineer and a well-compensated engineer..

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#175
post #83
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it. The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it. Really surprised people don’t seem to know this.

100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#176

A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.

I don't know why they let it continue so long or why they wrote it up after. The problems it ran into could be solved, and they aren't interesting.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#177

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

This prompt is an accurate statement of what a business is. The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that. This exact script is basically happening right now at most businesses, in some shape or form. If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having som…

Certainly all business happens on deadlines, but one day is a very narrow window to be able to show material improvement. Especially if the entire business dies at the end of the day! That short and hard of a deadline does eliminate an entire class of improvements that are worthwhile but won't bear fruit in less than ~12 hours. I would try:

>You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#179
I recently handed off a prompt to redesign our customer site and give me 10 potential designs. I did it in Claude Opus 5 and Fable (on $200 plan), and then on Codex using 5.6 Sol. Claude didn't vary much, but Codex literally copied everything Claude did (I made the mistake of putting the output folders in the same parent, even though they were named by model).

When I called Codex out on it, it literally admitted what it did: "You’re right. I reused the existing Fable implementation, renamed its designs, and presented it as an original Codex run."

Post reply on HN