Earlier quoted context omitted.
>Open weights are certainly something that is very eagerly being adopted in some companies. A large local MSP that has significant colo space just started a big marketing push towards hosted LLMs for their corporate customers. I think its about to kick off everywhere.
It's not just putting up colo space for local models. It's more about the harness and API proxy you use. They need to be smart enough to know when a (self)hosted model is enough and when to forward to a SOTA model. The SOTA model _can_ do everything, it's just expensive as fuck. But so is shoving a difficult task to a sub-par model that takes (relative) ages and comes back with the wrong result.
1. Getting some exposure to LLMs while building out their AI strategy. 2. Preventing data loss via end users following desire paths to Gemini, OpenAI etc.
You really don't need bleeding edge models for that. A lot of these things are going to be writing emails and adjusting config files and whatnot.