Live data from Hacker News

We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

arxiv.org

11–20 of 129 posts

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#11

> Native Persian speakers establish the human ceiling. Native speakers achieved an average accuracy of 81.8% on taarof-expected scenarios, demonstrating high but not perfect agreement. This establishes an appropriate ceiling for model performance and further validates our annotation approach I'm surprised human benchmark is that low. The canonical example of taarof, one I've seen elsewhere, is of a taxi driver insist…

> As an aside, there are elements of this sort of thing in Bay Area tech culture too.

The more general case of misaligned strength of a statement is widespread in other cultures as well.

E.g. I'm Norwegian, and it's not unusuals for Norwegians to use similarly soft language, though it's by no means universal. A statement like "perhaps it would be worth thinking about doing X?" will often mean "do X" or "do X right now!", and where you lie on the range from the literal meaning on one end and a direct order with an implied threat on the other extreme, may hinge on subtleties of intonation, which words are emphasised, and/or the personality and your relationship with the other person.

I live in the UK now, and my impression is that the same is true here but to a much lesser extent, and will then often be phrased in ways that may be easier to recognise by either being overly formal and/or wrapped in a layer of sarcasm.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#12

I'll stubbornly resist, and consider this a form of unnecessary protocol overhead, leading to even more shmancy sycophancy, which I do not fancy!

It's no different than GPT answering a prompt with "That's a wonderful idea!", except it's in a different language than English. It's a good thing if LLMs can do this in every language and for any culture with no compromise.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#13

> Native Persian speakers establish the human ceiling. Native speakers achieved an average accuracy of 81.8% on taarof-expected scenarios, demonstrating high but not perfect agreement. This establishes an appropriate ceiling for model performance and further validates our annotation approach I'm surprised human benchmark is that low. The canonical example of taarof, one I've seen elsewhere, is of a taxi driver insist…

I suspect that a lot of this form of cultural subtlety is designed to be hard on purpose. This allows individuals to show off and spar on a fairly harmless linguistic level - "he is so suave, she conveys her meaning just so" etc. Effectively it is a way to show off verbal + emotional intelligence in a way that doesn't look like showing off, since it's all about politeness.

The fact that it makes it hard for foreigners to productively engage is just a side benefit of the arrangement.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#15

> Native Persian speakers establish the human ceiling. Native speakers achieved an average accuracy of 81.8% on taarof-expected scenarios, demonstrating high but not perfect agreement. This establishes an appropriate ceiling for model performance and further validates our annotation approach I'm surprised human benchmark is that low. The canonical example of taarof, one I've seen elsewhere, is of a taxi driver insist…

I suspect that a lot of this form of cultural subtlety is designed to be hard on purpose. This allows individuals to show off and spar on a fairly harmless linguistic level - "he is so suave, she conveys her meaning just so" etc. Effectively it is a way to show off verbal + emotional intelligence in a way that doesn't look like showing off, since it's all about politeness. The fact that it makes it hard for foreigner…

My take on “Hard on purpose” is that it’s plausible deniability. It’s not for showing off, it’s to make the situation polite in a way that you can’t rationally attribute malice.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#16
I'm not sure what to think of this, on first impression I don't like it, but maybe I have a misguided impression on how this works.

Does ChatGPT properly handle western social customs? I'd say yes, and I presume that's because it has a truckload of data involving such customs and even some that explicitly talks about those customs. People do stuff, it gets recorded, then into the LLM.

In this case though we are talking about "artificially" generating content such that the LLM responds how the group making the content wants. Maybe that's something that was already done and I don't really have any ground to stand on?

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#17

I’m half Persian, and am relatively immersed in middle eastern culture still, but I sincerely wonder how I would perform on the benchmark too!

> ... is a sophisticated system of ritual politeness that emphasizes deference, modesty, and indirectness.

I'm Irish and think we have a similar culture of indirectness and politeness...

In the countryside anyway we're rarely very blunt... everything indirect...

  "You'll have a cup of tea Mary?"
  "Ah no.. sure I'm only after a drop"
  "Ah go on... you will"
  "Not at all, I'm grand"
  "Go on, go on, go on you will" etc (as in Father Ted)
I'm middle-aged now so maybe this has changed with the younger generation...

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#18
post #16

I'm not sure what to think of this, on first impression I don't like it, but maybe I have a misguided impression on how this works. Does ChatGPT properly handle western social customs? I'd say yes, and I presume that's because it has a truckload of data involving such customs and even some that explicitly talks about those customs. People do stuff, it gets recorded, then into the LLM. In this case though we are talki…

> it has a truckload of data

It would be "artificial" only if LLMs performed badly despite having an equal amount of data containing examples of eastern customs in its training set. Even that's arguable since we don't (didn't) have the benchmarks for this particular case before.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#19
post #7

I'll stubbornly resist, and consider this a form of unnecessary protocol overhead, leading to even more shmancy sycophancy, which I do not fancy!

That's kind of the point of politeness rituals in the first place, isn't it? To see who can be bothered to spend some extra effort to fit in and who doesn't care enough about the tribe to make the effort.

After having close contact with someone who do not follow these protocols, I came to conclusion they are much more then that. If you consistently do not follow them, it ends up being impossible to include you. Even if others put conscious effort into including such person, the end result is exclusion and everybody spending huge amount of extra effort.

Re: We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

#20
post #11

> Native Persian speakers establish the human ceiling. Native speakers achieved an average accuracy of 81.8% on taarof-expected scenarios, demonstrating high but not perfect agreement. This establishes an appropriate ceiling for model performance and further validates our annotation approach I'm surprised human benchmark is that low. The canonical example of taarof, one I've seen elsewhere, is of a taxi driver insist…

> As an aside, there are elements of this sort of thing in Bay Area tech culture too. The more general case of misaligned strength of a statement is widespread in other cultures as well. E.g. I'm Norwegian, and it's not unusuals for Norwegians to use similarly soft language, though it's by no means universal. A statement like "perhaps it would be worth thinking about doing X?" will often mean "do X" or "do X right no…

Okay, here’s what I’m wondering: How do you urge people to do discoveries or try things out when you’re reviewing something?

e.g. Do you think it would be better if we used a queue system here? Oh, no, I can try it but I had issues with blah blah etc.

Post reply on HN