You'd be surprised. Not enough teams still have their own evals. India[1] and the US[2] is my experience. To get many of them to understand the benefit of putting a couple of people to do data labelling and write a couple of verifiers for just a few days every few months was so difficult. Many folks have understood it all wrong and made allotments like "big model for this task" "small model for this task". Small was sometimes parameters, sometimes brand version number, or sometimes because it has "mini" in its name.
Proper LLM adoption beyond fucking around with claude code or github copilot is so low and is only going to go up as people figure it out. I also think the new cloud agents thing might accelerate adoption among these companies, since it's a bit more plug and play. But they are more likely to be stingy about it, so I'm not sure about the high margin expectations of certain model families.
[1] not witch, but BFSI. Surprisingly parts of witch companies have it down to science already, they embedded openai or anthropic or startups like devrev more than a year back.
[2] again not high fly SF companies, BFSI.