This kind of stuff just makes me think nobody has any clue how these things work. Why do I need a system prompt at all? Why do I need another black box AIs to review the code of the black box AI why can’t these things get code right the first time. Why is the best “coding model” in the world still making up APIs that don’t exist and do seemingly random unreleased changes that it wasn’t prompted for. Why do these mode…
2. Meta tokensations of LLMs. Imagine entire output of the LLM as just a single token, so they inspect previous one.
3. Non-deterministic output.
4. Benchmaxxing. Mobile phones are doing the same thing.