Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

91–100 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#91
post #54

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.

> Results that arrive after the deadline do not exist

Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#92

> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt. At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using too…

[deleted]

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#93
post #83
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it. The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it. Really surprised people don’t seem to know this.

I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”.

If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#94

Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack. That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.

The AI paperclip case, however, is coming on extraordinarily strong.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#95

Earlier quoted context omitted.

Kinda some kettel logic here no? Is it not rigorous enough, or is it in-line with typical failure rates?

Rigor would be trying it more times so that you can perform statistical tests against some established baseline rate. Feasibility without funding would be the problem, as alluded to in another comment.

I am just trying to (gently) suggest you did not frame your points here in a good or convincing way, but thanks for the explanations here anyway.

Sure sounds like there would be a lot to think about either way!

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#96
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

I feel like new graduates will need to start taking linguistics, psychology and public speaking classes in order to understand why and how subtext matters, and how to control it. Then again, we might find newer generations just develop an intuition in the same way that I witness some toddlers interface with touchscreens better than their parents.

[deleted]

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#97
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The followi…

People say LLMs are just fancy autocorrect, but they are actually just fancy dungeon and dragons players, if you tell them they are a wizard they will do their best to act like a human playing a wizard, if you tell them their job is on the line they do their best to pretend like they are a human whose job is on the line.

It's all just roleplay.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#99

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

This prompt is an accurate statement of what a business is.

The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that.

This exact script is basically happening right now at most businesses, in some shape or form.

If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business.

Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies.

How did you expect the prompt to be written?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#100

A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.

Was that the one that gave away PS5s?

Yep!
Post reply on HN