> 2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc.
The problem with this angle is that it is still absolutely terrible at doing call center/customer service work, and the profitability story is that the price is going to go up rather than go down.
For repetitive physical labor in a controlled environment I'm slightly more bullish, but if you control the environment, you mostly don't need AI. You just use traditional deterministic methods, and send a person in when things get stuck or things are by nature irregular.
> 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc.
Those who can accept failure cheaply can't necessarily detect failure cheaply. A ton of insane attempts will have to be picked through carefully to find the candidates for success, because the lack of a thought process makes AI bad in random, inhuman ways. This is basically a version of 3) that wishes away tests. It will be (and is) certainly helpful to replace interns and aid in rapid prototyping, but not because failure can be accepted, but because those are things that are tightly supervised. According to the world thus far, that is resulting in anything from -15% to +25% productivity gains. I'm not seeing it as a game changer simply because if it was, I'd expect to have seen a lot more useful, original software products by now and I haven't. I've just seen old ones get buggier or rewritten in Rust.
I'm only buying 3): when you just want a machine to randomly enumerate through a search space looking for things that make the carefully constructed tests pass. That's a very good thing, though. But as you say, it's not a game changer because you still have to write the tests.
but
> 5. navier-stokes and statements in pure mathematics like it are the absolute best case scenario for agentic work against rigorous specification. the theorem statement itself is already a rigorous specification. it has undergone decades of auditing by the mathematical community and its rendering in lean is a straightforward translation defined in terms of battle-tested mathematical objects from mathlib. the verifier, the lean theorem prover, has been extensively audited and specifically designed to avoid the types of unsoundness that would make it vulnerable to reward hacks. even lean and theorem provers like it are not invulnerable: soundness bugs have allowed LLMs to launder bogus proofs through the proof kernel before and it is not improbable that more such bugs exist. this is the rosiest setup; the vast majority of human knowledge work does not look like this. i'll comment below on the few areas of knowledge work that do resemble pure mathematics in this respect.
This is the real deep point, and one I've been repeating since I heard Navier-Stokes was a fraud.
This is exactly where I expected that LLMs would do well, and they are not.
It shows that I have a basic misunderstanding of LLMs, and that misunderstanding is causing me to think that they have more potential than they have actually shown.
Maybe the nature of the architecture, where it picks out features, intrinsically limits its ability to search a solution space?
Maybe the fact that they modally predict what someone might say, and nobody has said a thing as of yet (when many people were knowledgeable enough to have, if it is correct), means that the LLM is not going to say it either?
Maybe the fact that it consumes all information and blends it in a structured way, instead of synthesizing an entire space from a relatively very small amount of input like a human does, means that it won't ever accidentally synthesize something that can't be pieced together from things that have already been said? Is its accuracy its flaw, where a human's "mistaken" synthesis might ultimately correct everyone's understanding?
Really not beating the charge of being a stochastic parrot. It might just be that we were underestimating stochastic parrots; if a million monkeys on a million typewriters were all getting treats when they satisfied a trainer who wanted to see a new work of Shakespeare; they could look at his published work, and they could watch each other type and when each other got treats; whenever they successfully spelled a word or put words into an intelligible phrase, that was made into a keyboard key for a group of sentence monkeys, and the successes of the sentence monkeys were made into keys for the paragraph monkeys, etc... could you get something that passed for mediocre, drunken Shakespeare in a thousand years? Or maybe even 10?