Earlier quoted context omitted.
> It checks these using an LLM which is instructed to score the user's prompt. You need to seriously reconsider your approach. Another (especially a generic) LLM is not the answer.
What solution would you recommend then?
I don't know what I would use, but this seems like a bad idea.