Earlier quoted context omitted.
To me, this doesn't show the weakness of current models, it shows the variability of prompts and the influence on responses. Because without the prompt it's hard to tell what influenced the outcome. I had this long discussion today with a co-worker about the merits of detailed queries with lots of guidance .md documents, vs just asking fairly open ended questions. Spelling out in great detail what you want, vs just g…
So which approach worked better?
I have a hunch if you asked which approach we took based on background, you'd think I was the one using the detailed prompt approach and him the vague.