>Also stop saying “please” to an LLM. It does not have any feelings.
It emulates having feelings and it's a next token predictor. The next token in a dataset where you respond to an engineer rudely and call them stupid is rarely said engineer locking in and delivering incredible code. You have to play along to get the output you want.
> Again, I recommend people try running small LLMs locally where temperature and other settings are fully exposed and configurable to see this themselves.
This is like saying you should experiment with a paper airplane to see why fighter jets are overrated. Also the general understanding of temperature is not super solid here - it's not just that its "too boring" without it - random sampling is required for the models to work.
>Why are benchmarks showing they're still improving?
Ironically this section is far too generous to LLMs and benchmarks. lLMs cheat and companies benchmax. Don't trust benchmarks. They lie
Also in the what I do section - are these using local LLMs too? "Sometimes looping llms on itself can make it fix its own errors" feels very 2025 - the modern state of things is more "we've given up on one shots and getting it to produce the right answer immediately, set it up with a test harness so it can fix its own mistakes and let it loop otherwise it won't work." Also "asked fellow developers to make sure their code is well structured, easy to follow and documented. LLMs unfortunately make it easier for people to cheat in this regard" - really? Easy to follow? Maybe gpt models with good steering but trying to get an anthropic model to speak coherently and clearly and write documentation that isn't incomprehensible slop is a Herculean effort