I've heard something along the lines of "Claude is like a compiler: source code is the new object code, you don't look at that anymore" many times. And I don't really think this is true. Compilers are usually deterministic, and whilst we can find edge cases, it's nothing like an AI agent writing all the code for you. I think you have two choices, given the Claude is a code generator and not a compiler: (a) you review…
Part of the value of LLMs and humans is nondeterminism. Pair nondeterministic output with strictly verified results (proper tests) and you can create a useful working system. Humans can't build a system perfectly, and agents definitely can't. And agents (like humans) will build a different system every time even with the same prompt. Even if it's just trivial differences like array vs linked-list, there are still dif…
What do you think happens in terms of testing/reviewing going forward? I'd really appreciate your thoughts on that.