> 1500 t/s [...] all you can eat tokens as fast as you can eat them
Pfft, you can eat tokens far far faster than that. Many orders of magnitude more. Just switch from frontier-style bespoke artisanal pets to cloud burst-parallel, latent subspace exploring/exploiting/searching, mass ensembles of cow herds.
There's N-Version Programming. Work the problem in English, in Chinese, in Haskell, Lisp, Rust, etc. Then work ports to the target lang.
There's design space sampling. Work the problem emphasizing performance, or security, or monitoring, readability, etc. Then work a synthesis.
There's non-determinism sampling. Work the problem order 10 or 100 times. Then work to combine the best bits from each.
There's sample synthesis. NP-hard aggregation of insights.
There's genetic exploration. Work populations of trees of work variants under selective pressure.
There's repo quantum superpositions of implementation space. The unspecified remains indeterminate - state space collapse occurs not upon each edit/commit, but as JIT-synthesized fuzzing/search upon each execution.
There's maintaining a pretty dev UI, but that >>10k tok/s is trivial, because like symbiotic adversary cocreation, fine-grain agent swarms, scenario analysis/forecasting, etc, etc, it is unlike the preceding items... which scale combinatorially.
"All you need is 1500 t/s"? "All you need is 640k RAM" is only 5 orders of magnitude off from 64 GB. It takes "All you need is a single Intel 3101's 64 bits", to get 9 orders of magnitude from 64 GB. Then datacenters...