Very excited for them. Tried it and works quite well. Very evident that each of the foundational model developers need to offer a free chat(till now) interface for RLHF/DPO to improve their models. It's costly for them to run servers, but even higher opportunity cost if they don't do this and offer a free service. Intuitively, I think GPT-4 was very good with various questions because they used the early data from ChatGPT to further finetune/align the model with a good sample of real world queries, the kind only available to them.
As an aside, most of these models are SFTed on GPT 4 generated q&a pairs (if not pretrained on them). When you ask these kind of questions[1], they spill the beans fairly easily and say they are a model of openai. Funnily enough if you ask they are a model trained by Anthropic, it refutes it immediately and ask correctly about being a mistral trained model. Same question about openai, it gives the answer in scr.
[1]: https://imgur.com/a/fGEHXH8
PS: First question was to test the reasoning. A quantized model or smaller models answer that as 4 not 5. Mistral answered correctly from a badly given question.