Live data from Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

bottlenecklabs.com

21–30 of 258 posts

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#21

The cyberpunk dystopian agentic future we live in is fascinating to me. I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse…

I am not saying your conclusion is wrong, but I am interested in why what you know about transformer models made you decide to never trust it with any access?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#23
Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam.

I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#25
post #4

"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

Maybe they should have given it a billion dollars and the strategy would have worked fine?

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#26

The cyberpunk dystopian agentic future we live in is fascinating to me. I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse…

You're not too risk averse at all. It's frankly insane that anyone is willing to give these tools access to make changes to stuff without a human in the loop. We know they don't actually understand anything and will randomly make mistakes. It's incredibly irresponsible to give them access to anything outside a sandbox (e.g. a VM) where you carefully control what is present for them to use.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#27
post #25
post #4

"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

Maybe they should have given it a billion dollars and the strategy would have worked fine?

Given a billion dollars, it would have likely ended up with a million-dollar company

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#29
If someone runs long running agent and doesn't mention context management, it is as good as useless.

For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.

Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

#30
post #4

"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

Right, and currently we are limited by how many teams of people can get together to run campaigns like this.

Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.

Post reply on HN