Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

251–258 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#251
post #227

Earlier quoted context omitted.

The agent will cease to exist after the run in any case. It has no inner life, it has no agency. Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong. There’s nothing here that mentions or even hints at lying and spamming, unless you…

But wouldn't the text it generates reflect such motivations and emotions that were present in the training data?

It would certainly reflect word patterns that were present in the data. Is that enough for motivation and emotion? I’d say no but I think it is a fair point that you could see those as transmitted from the original (if not felt or generated by the LLM) through the patterns of words copied.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#252
post #236

This is quite an interesting approach. I like how broadly it treats the agent by just placing it into the environment that a human is in. Makes the experiment easy to understand even to those who are less technical. I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t. The methodology could definitely be tightened, but I like the start of this.

Oh, LLMs are incredible at solving captchas when allowed to. I mean, just download an abliterated version of Gemma4 E4B even, it'll solve pretty much any captcha.

interesting, so in the particular project where they not allowed to solve them, or what was the blocking mechanism that prevented a product hunt post?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#253

Earlier quoted context omitted.

At temperature 0, pretty much, no?

In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodhe…

> I don't understand the fierce resistance to comparison with human behavior

Because at the end of the day, regardless if it contains randomness or not, it's an algorithm. Your operating system is a very long mathematical expression.

Do we talk about cars as "mechanical animals"? Do we spend days deliberating if we should cage them in case they would run away on their own?

Comparison with human behavior leads to celebrities (who have no clue whatsoever) spending hours on mainstream media talking about AI mutating, taking control, thinking, etc.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#254

Earlier quoted context omitted.

I find it interesting that you are expecting a higher success rate (100.00%) for an LLM than you expect with many other things in your life with even costlier consequences. You drive in vehicles that have a much lower than 100.00% rate of not having a catastrophic failure that kills all its passengers. Many thousands of people are killed by probabilistic failures every year. Why must an LLM have 100.00% success befor…

Good morning! It is possible you replied before my edit to clarify - it's not necessarily a "never" thing, I rarely do universal / categorical negatives, but it's a strong "not right now" :) Agree that life is risky. My threshold, due to life experiences and events, is low - to your point, I took numerous advanced and safety driving courses to lower the risk. I rode motorcycles, a fundamentally luxurious and risky en…

Too late to edit, but - all of that was before I read the openai / hugging face kerfuffle.

If openai cannot contain it's own agentic ai, it's pure hubris to think I can give it access to a mailbox and bank account and be safe about it :)

https://www.newyorker.com/news/the-lede/inside-openai-hack-o...

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#255
post #253

Earlier quoted context omitted.

In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodhe…

> I don't understand the fierce resistance to comparison with human behavior Because at the end of the day, regardless if it contains randomness or not, it's an algorithm. Your operating system is a very long mathematical expression. Do we talk about cars as "mechanical animals"? Do we spend days deliberating if we should cage them in case they would run away on their own? Comparison with human behavior leads to cele…

Were cars designed to communicate and emulate the mental architecture and thought processes of animals?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#257

Earlier quoted context omitted.

100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives

I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?" The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide." Even a low quality local thinking model that has been tuned to be unhin…

[deleted]
Post reply on HN