Earlier quoted context omitted.
My gut feel is Anthropic is very technical and pedantic which makes their models really technical and pedantic. They're top at code and technical benchmarks but anecdotally I've found OpenAI to be significantly farther ahead for general usage. Opus 4.8 will burn 10k tokens trying to answer something 100% whereas GPT-5.5 will burn 2k getting it 90% which is good enough for many things. Some personal testing on a "help…
The problem is that the remaining 10% can bite you in bad ways. I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff. When planning the trip, my sister used ChatGPT, which helpfull…
Why trust an LLM with information like bus schedules? They fuck up things like this routinely.