Earlier quoted context omitted.
What about OpenAI Structured Outputs? This seems to do exactly this.
I'm building this type of functionality on top of Llama models if you're interested: https://docs.mixlayer.com/examples/json-output
What are you doing underneath, here? If thats secret sauce, I'm curious what you're seeing in tokens/sec on ex. a phone vs. MacBook M-series.
Or are you deploying on servers?