At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.
We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
31–40 of 258 posts
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#32Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#33If someone runs long running agent and doesn't mention context management, it is as good as useless. For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.
The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#34"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#35The article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer) Also what is the failure rate of tech businesses again? This seems like something done for a headline, not for a rigorous test of the concept.
Kinda some kettel logic here no? Is it not rigorous enough, or is it in-line with typical failure rates?
Re: We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
#36Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam. I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
That's because it's an advert, not an experiment