Earlier quoted context omitted.
Lots of interesting issues: - The agent has a tool to set it's task to 'completed', 'failed', or 'needs_help', with the last one being a option for human in the loop scenarios. Sometimes the agent gets lazy and says it needs help prematurely. - Additionally, the agent can create subtasks for itself, either to run immediately, or to schedule in the future. Here it again can call that tool a bit too eagerly, filling du…
> showing flashes of brilliance A “flash” of anything is also called a fluke, or a coincidence. The dumbest moron can have a flash of brilliance on occasion. So could a random word masher. Consistency is what matters. > and we're gaining more and more conviction that this is the right form factor Are we? Who’s “we”? Because it looks to me like the LLM approach is lacklustre if you care about truth and correctness (wh…
"We" are a YC-backed startup: https://www.ycombinator.com/companies/bytebot.
Re: truth and correctness, their are different tolerances depending on the type of task.