> Native Persian speakers establish the human ceiling. Native speakers achieved an average accuracy of 81.8% on taarof-expected scenarios, demonstrating high but not perfect agreement. This establishes an appropriate ceiling for model performance and further validates our annotation approach I'm surprised human benchmark is that low. The canonical example of taarof, one I've seen elsewhere, is of a taxi driver insist…
The more general case of misaligned strength of a statement is widespread in other cultures as well.
E.g. I'm Norwegian, and it's not unusuals for Norwegians to use similarly soft language, though it's by no means universal. A statement like "perhaps it would be worth thinking about doing X?" will often mean "do X" or "do X right now!", and where you lie on the range from the literal meaning on one end and a direct order with an implied threat on the other extreme, may hinge on subtleties of intonation, which words are emphasised, and/or the personality and your relationship with the other person.
I live in the UK now, and my impression is that the same is true here but to a much lesser extent, and will then often be phrased in ways that may be easier to recognise by either being overly formal and/or wrapped in a layer of sarcasm.