I'm doing a small local version of this test to pull moods and themes out of song lyrics. I've found that one prompt, 1 model, 11 runs even gives different results.
There's no consistency over multiple runs of the same prompt on the same model.
Also, for the purpose of music lyric analysis Qwen 4B is laughably bad. Like it's going out of its way to be extremely wrong, misunderstand the prompt. When it does correctly understand what I asked for it ALWAYS tells me that the mood is Angry. Sometimesiit just gives back all the lyrics. Sometimes it claims that it doesn't have the list of moods or the lyrics and tells me I should look them up on the Internet first.
All with the same prompt every time.
Are models being overtuned for coding tasks?