Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

71–80 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#72
I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, learn, try something else, repeat.

It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#73
post #67
post #60

Earlier quoted context omitted.

Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.

I agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating. Granted, this can probably be tuned for.

And really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#76
post #32

That's better than the performance of the average new hire. 24 hours to push a product with a very narrow market is not much.

If a newly hired colleague lied like this I would strongly argue to my immediate superior to end their probation period/employment immediately.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#77

A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.

Was that the one that gave away PS5s?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#78
This is dumb. You need two teams ideally the same app or business in different markets for a business quarter.

One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#79
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The followi…

And yet they're trained on the corpus of human writing. They may not act like humans but they do act like human writing.

"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#80
post #65

Earlier quoted context omitted.

Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The followi…

You can literally read their thoughts if you run an open model, they look like pretty human thoughts to me, albeit a neurotic human.
Post reply on HN