Earlier quoted context omitted.
How? The agent does something incredibly difficult, and there is no human code review. Why do you think it doesn't match ?
Literally every top 10 LLM can tell you why. Maybe we’re in „this is the end of making other people spell it out for you” era?
The fact is, the agent did something amazingly difficult, that zero human on earth could do. Read the spec if you believe someone could do it. It did it clean, everything is public, and it works perfectly. Zero code review.
https://aisovereignlabs.ai/docs/case-study/liveSession/logs/...
99.9% of the IT dev tickets are easier than this, excluding off topics like discrete math, computer vision algorithms, etc, which are irrelevant here.
Do you think there's anything you've done, with human code review, in the past 30 years, matching the difficulty of this ?