Not enterprise, but one thing that transferred from running a lot of agents on a small budget: the cost that got away from us was never the big deliberate runs, it was the default behaviours nobody had looked at. Two we measured and turned off:
1. Re-reading whole files after every edit "to verify". The edit result already tells you it applied. Forbidding the re-read unless a test fails cut a noticeable slice of tokens on long sessions with zero quality change.
2. Guessing loops. An agent that gets a fix wrong twice will happily try a third and fourth variant. We put a rule in the harness: after the second failed attempt it has to add instrumentation and report what it observed before it is allowed to change code again. That turned several hour-long loops into ten-minute fixes, and the token savings were incidental to the time savings.
On limits: a hard monthly cap per person mostly moved the spend to the last week of the month. What worked better for us was a visible remaining-budget gauge in the tool people actually work in, so the number is in front of them while they decide whether to kick off another run. People self-regulate surprisingly well when the meter is on the dashboard rather than in a monthly report.
If you do tier by role, I would tier by "how expensive is a wrong answer" rather than tech vs non-tech. A non-technical person running one careful summarisation a day is cheap; a developer with an agent in a retry loop is where the money goes.