Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.
I am shocked that it speaks the way it does when it was trained on random stuff it doesn’t have rights to. They say they trained it on databases they had bought access to etc. And it seems that way. Because how does ChatGPT: 1. Do what you ask instead of continuing your instructions? 2. Use such nice and helpful language as opposed to just random average of what people say? 3. And most of all — how does it have a str…
And by curating your sources you are of course going to help the model to achieve something a bit more sensible as well. Finally: you are probably not looking at just one model, but at a set of models.