That solution actually makes great sense. So Apple won in some strange way again? Guess there are limitations on size of the models, but if top-tier models will getting democratized I don’t see a reason not to use this API. The only thing that comes to me is data privacy concerns. I think batch-evals for non-sensitive data has great PMF here.
Because they were already at the finish line with Apple Silicon.
> I don’t see a reason not to use this API. The only thing that comes to me is data privacy concerns.
The whole inference is end-to-end encrypted so none of the nodes can see the prompts or the messages.