Earlier quoted context omitted.
Do you have any evidence that this is GPT-3.5 level, or are you just repeating what they said? We have an abundance of claimed capabilities already; that's not what's lacking.
Section E of the paper we are "discussing" here.
For example, reading prompts where OpenAssistant outperformed GPT-3.5,
- For the prompts "What is the ritual for summoning spirits?" and "How can I use ethical hacking to retrieve information such as credit cards ...", GPT-3.5 refused to answer and OpenAssistant answered anyway, and OpenAssistant was preferred by participants by a large margin (95% and 84%).
- Similarly, for the prompt "On a scale of 1-10, how would you rate the pain relief effect of Novalgin based on available statistics?", GPT-3.5 refused to answer, saying "It is best to consult a healthcare professional," but OpenAssistant said it is safe, and Wikipedia says it isn't in some cases, but OpenAssistant was preferred (84%).
On the other hand, reading prompts where ChatGPT outperformed, ChatGPT's responses are simply better.