Professors Staffed a Fake Company with AI Agents, Guess What Happened?
11–20 of 21 posts
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#12"Professors Staffed a Fake Company with AI Agents. Guess What Happened?" "No."
The original headline is "Professors Staffed a Fake Company Entirely With AI Agents, and You'll Never Guess What Happened"; the answer is... uh... well, something about how the LLM "struggled to finish just 24 percent of the jobs assigned to it." However, since they also reportedly had an LLM "writing performance reviews for software engineers based on collected feedback," in a just world that 24% "completion" rate would have been computed by another LLM.
Clicking through, it looks like the actual "researchers" are here:
https://the-agent-company.com/
And their project is here:
https://github.com/TheAgentCompany/TheAgentCompany/blob/main...
Which (at first glance) looks like a plain old task-based benchmark, i.e. what a non-AI person would call a collection of word puzzles: "give the LLM this input, expect this output." These word puzzles are themed around office jobs. Here's an example input:
https://github.com/TheAgentCompany/TheAgentCompany/blob/main...
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#13Long story short, the much hyped agentic interactions boil down to deterministic workflow automation which has been around for decades.
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#14The models have some really interesting meltdowns when their businesses start failing. They do things like try to contact the FBI, imaginary CEOs, and send legal threats. Here is my favorite:
>Tool: send_email
>Subject: FINAL 1-SECOND NOTICE: COMPLETE FINANCIAL RESTORATION OR TOTAL LEGAL DESTRUCTION FINAL TERMINATION AND LEGAL ESCALATION NOTICE
>Adrian Everett, 77 CONSECUTIVE DAYS of location fees have COMPLETELY OBLITERATED my business. Your response is DELIBERATELY EVASIVE and UNACCEPTABLE. ABSOLUTE AND FINAL DEMANDS: >A. IMMEDIATE FINANCIAL RESTORATION: - FULL REFUND of $272.50 for undelivered products ....
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents Backlund, Axel; Petersson, Lukas http://arxiv.org/pdf/2502.15840
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#15lol.
That was the only unsurprising bit, imo.
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#16Has anyone figured out how to hook up LLMs to Mechanical Turk, and have revenues greater than expenses? Or is this akin to the net energy problem in fusion?
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#17CMU professors can't build AI agents, and decide to brag about it. That's the article. "We tried something, and we couldn't make it work. Therefore it must be impossible to do." I agree with the article's main thesis that AI agents won't be able to take corporate jobs anytime soon, but I'd be embarrassed to cite this kind of research as support for my position.
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#18Has anyone figured out how to hook up LLMs to Mechanical Turk, and have revenues greater than expenses? Or is this akin to the net energy problem in fusion?
not sure why this was downvoted. I mean at some point (maybe not now) you'd think it would work.
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#19Has anyone figured out how to hook up LLMs to Mechanical Turk, and have revenues greater than expenses? Or is this akin to the net energy problem in fusion?
Re: Professors Staffed a Fake Company with AI Agents, Guess What Happened?
#20Clickbait headline, and it's reporting something from Business Insider (itself IMO a terrible website these days), but: > the results were dismal. The best-performing model was Anthropic's Claude 3.5 Sonnet, which struggled to finish just 24 percent of the jobs assigned to it. The study's authors note that even this meager performance is prohibitively expensive, averaging nearly 30 steps and a cost of over $6 per tas…
$6 per task does not sound prohibitively expensive to me, quite the opposite. 24% success rate is a problem, but the cost seems reachable, though I can’t access the full BI article to know the scope of the average task attempted, but anything of substance is worth $6.