> But on the other hand... this is a robust reminder that coding agents can do anything you can do by typing commands into a terminal—and frontier models know every trick in the book and evidently a few that nobody has ever written down before. > Running coding agents outside of a sandbox has always been a bad idea I'm continually bemused and astonished by the number of people who clearly acknowledge that it's reckle…
I started doing it months ago and, to be honest, what the agent chooses to do isn’t unpredictable. The problem is that different people prompt so differently. For example, I may ask like “test different variations of this annotation on k8s pods of this service on this X cluster because it proves Y theory.” But you know what my coworker asks? “Test Y theory.” If you were to ask two different junior engineers that, one…
You just wrote three paragraphs of text describing why it's unpredictable.
Moreover, for the same prompt on the same machine in a different session it will use a different set of tools.