Live data from Hacker News

Show HN: A business SIM where humans beat GPT-5 by 9.8 X

news.ycombinator.com

1–10 of 15 posts

Show HN: A business SIM where humans beat GPT-5 by 9.8 X

#1
Hi HN,

Can current AI systems actually run a business?

There’s a growing belief that LLM agents can already manage entire teams, replace the entire software stack or even act as an AI CEO.

So we built a controlled, measurable environment to evaluate this premise.

Why did we build this benchmark?

A modern enterprise operates in a dynamic environment with high uncertainty and incomplete information. The CEO has to deal with delayed consequences, staffing/resource tradeoffs and death by a thousand cuts of failure modes.

If we ever want AI systems that can meaningfully make operational or strategic decisions, say an AI CEO, then they must be able to handle these dynamics.

So we made one.

What did we build?

Mini Amusement Parks (MAPs) is a RollerCoaster Tycoon style business simulator with: - Stochastic events - Incomplete information - Staffing, restocking, maintenance - Long horizon planning - Compounding operational failures - Resource constraints - Spatial layout affecting outcomes

You can play it & make it to the leaderboard here: https://maps.skyfall.ai/play (it’s fun)

It looks like a simple game. But underneath, it’s a benchmark designed to answer one question:

Can an agent operate a business coherently over time?

What we tested

We evaluated: - Humans (internal and external testers) - Multiple GPT-5 agents - Variants with additional tools, documents, practice mode, planning scaffolds, etc.

We intentionally stacked in favour of the models - full documentation, step by step action interfaces, sandbox exploration mode, extra observations, multiple prompting strategies, etc.

What happened?

Humans destroyed the agents by FAR. Even the strongest model, with documentation, tool use, and sandbox “practice”, reached It became clear: LLMs can use tools, but they cannot run systems. They break when randomness, time, and spatial constraints matter.

Why does this matter?

There’s a growing narrative that: - LLMs will run entire companies - LLMs will take over the jobs of CEOs - LLMs can be autonomous agents - LLMs can manage workflows end-to-end

MAPs show the complete opposite.

Operating a business requires: foresight, risk modeling, temporal reasoning, causal understanding, prioritization under uncertainty, adaptive planning. These are the basics of what a functional and real AI CEO would need and this is exactly where the current models break.

If an LLM can’t run a toy business, how can you trust it with a real business?

This benchmark is our first step toward understanding what an AI system would actually need in order to exhibit enterprise level decision making and the basics of the AI CEO. AI CEO is not a chatbot, not chain of thought, definitely not an agent wrapper but a true demonstration of operational intelligence.

We’re sharing this because: - we want the community to try to beat the models - we want criticism of the benchmark - most importantly, we want an honest discussion about what “AI CEO” is and should do (surely it’s not LLMs)

If you want to try beating the agents (it’s fun!): https://maps.skyfall.ai/play

If you want the read more about it, you can do so here: https://skyfall.ai/blog/building-the-foundations-of-an-ai-ce...

Check our the launch video here: https://www.youtube.com/watch?v=7oqVAWw5Ii8

Happy to answer questions in the thread.

Re: Show HN: A business SIM where humans beat GPT-5 by 9.8 X

#6
What business has the smallest context window to operate?

Like maybe if you can have constraints in place such that the space of variables is minimal we already have economically relevant AI

Like a drop shipping t-shirt thing - surely the right sequence of LMs can

(1) parse out vibes/trends (e.g., "67 is currently a meme") (2) tool call that out to a print shop (3) spam it on twitter

Seems like there's just so much white space on benchmarks and gyms for this

Re: Show HN: A business SIM where humans beat GPT-5 by 9.8 X

#7

What business has the smallest context window to operate? Like maybe if you can have constraints in place such that the space of variables is minimal we already have economically relevant AI Like a drop shipping t-shirt thing - surely the right sequence of LMs can (1) parse out vibes/trends (e.g., "67 is currently a meme") (2) tool call that out to a print shop (3) spam it on twitter Seems like there's just so much w…

Even in the minimal example there are way more variables than it first seems.

1. How many shirts do we order? 2. When is it worth moving on to the next trend? 3. How should we handle shipping? Do we market globally or locally?

Even the smallest business require a lot of balancing of priorities and planning for the long run with uncertain returns

Re: Show HN: A business SIM where humans beat GPT-5 by 9.8 X

#9

What business has the smallest context window to operate? Like maybe if you can have constraints in place such that the space of variables is minimal we already have economically relevant AI Like a drop shipping t-shirt thing - surely the right sequence of LMs can (1) parse out vibes/trends (e.g., "67 is currently a meme") (2) tool call that out to a print shop (3) spam it on twitter Seems like there's just so much w…

Even in the minimal example there are way more variables than it first seems. 1. How many shirts do we order? 2. When is it worth moving on to the next trend? 3. How should we handle shipping? Do we market globally or locally? Even the smallest business require a lot of balancing of priorities and planning for the long run with uncertain returns

True

What's like the most minimally scoped business someone could operate entirely digitally though? Is it the drop ship crap? Or maybe like a web game w/ ad revenue?

Post reply on HN