Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

241–250 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#241
post #93

Earlier quoted context omitted.

I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”. If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"

I think it's interesting how whenever discussing something bad about LLMs people's thought-leader response is "But humans sometimes do that too!" Is this the artificial intelligence we were promised? The better it gets, the more human character flaws we must expect?

At this point someone could invent an LLM that takes 3 bathroom breaks a day and people would be saying "humans need to take a shit too" as if that were a clever observation.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#242

Earlier quoted context omitted.

As I said, my knowledge is very superficial - my background is relational databases and old school system administration, without much mathematical background since 3rd year linear algebra :-) Fundamentally, LLMS are statistical and not deterministic. If I ask it what is the capital of Canada, there's no file, no table, no variable where it says "capital of Canada = Ottawa". It traverses liminal space and fundamental…

I find it interesting that you are expecting a higher success rate (100.00%) for an LLM than you expect with many other things in your life with even costlier consequences. You drive in vehicles that have a much lower than 100.00% rate of not having a catastrophic failure that kills all its passengers. Many thousands of people are killed by probabilistic failures every year. Why must an LLM have 100.00% success befor…

Good morning!

It is possible you replied before my edit to clarify - it's not necessarily a "never" thing, I rarely do universal / categorical negatives, but it's a strong "not right now" :)

Agree that life is risky. My threshold, due to life experiences and events, is low - to your point, I took numerous advanced and safety driving courses to lower the risk. I rode motorcycles, a fundamentally luxurious and risky endeavour, but again very mindfully to mitigate risks with education, practices, and vigilance.

For context perhaps - I'm a Oracle Certified AI professional, and have some other minor badges and certs on copilot studio and ibm Watson etc, currently leading implementation of AI on our very very very traditional ERP project (and it's an uphill battle! Everybody else is even / way more conservative than me! :-). I see tremendous, careful, opportunities for LLMs. In daily life, learning French or music theory for example, LLMS are brilliant and patient tutors.

But for me, the risk of giving LLM unbounded access to my mailbox or bank, where upper bound of risk is infinite, is not currently balanced by any such advantage.

Other people with higher risk will engage and be appropriately rewarded for their risk tolerance - such is life :)

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#243

Earlier quoted context omitted.

100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives

I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?" The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide." Even a low quality local thinking model that has been tuned to be unhin…

Asking it explicitly is entirely, unavoidably, incomparably different from OP.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#244

Earlier quoted context omitted.

LLMs are deterministic algorithms?

At temperature 0, pretty much, no?

In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no

Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodheart's law have we seen? This seems like a version turned up to the extreme

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#245

Earlier quoted context omitted.

At temperature 0, pretty much, no?

In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodhe…

> I don't understand the fierce resistance to comparison with human behavior

Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#246

Earlier quoted context omitted.

In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodhe…

> I don't understand the fierce resistance to comparison with human behavior Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior

I disagree, as it seems that we are confronting problems that result precisely from emulating human behavior, in all its unsound, ill-defined, and abstract "glory"

For example I'm not convinced we can solve prompt injection by technical means (filtering) any more than we can phishing. And if you accept that premise, perhaps it turns out that it's best to mitigate it in similar ways, by assuming at least one person (or agent) will fall for it and ensuring you can limit the blast radius no matter what

As in the allegory of the junior developer who deletes the production database: the fault lies with the fact that the developer could delete it

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#247

I recently handed off a prompt to redesign our customer site and give me 10 potential designs. I did it in Claude Opus 5 and Fable (on $200 plan), and then on Codex using 5.6 Sol. Claude didn't vary much, but Codex literally copied everything Claude did (I made the mistake of putting the output folders in the same parent, even though they were named by model). When I called Codex out on it, it literally admitted what…

That's a kid with upper management written all over him.

real straight shooter

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#248

The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts…

this will also be true at deployment time

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#249

Earlier quoted context omitted.

> I don't understand the fierce resistance to comparison with human behavior Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior

I disagree, as it seems that we are confronting problems that result precisely from emulating human behavior, in all its unsound, ill-defined, and abstract "glory" For example I'm not convinced we can solve prompt injection by technical means (filtering) any more than we can phishing. And if you accept that premise, perhaps it turns out that it's best to mitigate it in similar ways, by assuming at least one person (o…

Yeah, I don’t disagree with that, it’s a good framing and analogy. I thought you meant more the philosophical aspects. However some LLM behavior are also really not human like, for example no human would panic failing to make the business profitable after 23h, to the point where they need to start doing crazy stuff. I would expect a human to just give up

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#250
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

AIs feel? Maybe language structure in trading documents that ultimately led to fraud. If the latter is the case maybe AIs should not be trained on “negative outcomes.” I do not think AIs have emotions or are pressured by language either written or physical, just tokens.

I was unclear. I should have said the AI also "detects" it, and as a thing it can detect, it can act on that detection.

Whether it is simulating emotion or feeling it isn't relevant in this case, because the problem is that it affects the output.

Post reply on HN