Thankfully archive.org exists, otherwise it would not be possible to get good training data in a few years when the internet is flooded with AI content.
Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM
Google denies training Bard on ChatGPT chats from ShareGPT
111–120 of 342 posts
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#112I don't care at all about this from a copyright or data ownership perspective, but I am a little skeptical that it's a good idea to be this incestuous with training data in the long run. It's one thing to do fine tuning or knowledge distillation for specialized domains or shrinking models. But if you're trying to train your own foundation model, is relying on output from other foundation models going to make them lea…
The whole thing is a plateaued feedback loop.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#113Earlier quoted context omitted.
Bard is only a week old and has a large "experimental" sticker on it. Besides its UI is better and the answers are succinct which I prefer.
They literally copied the Chatgpt UI, lol, only it looks like a dated Google UI. How do you prefer answers with less data?... that's crazy.
As for the UI it's a take on the tried and true chat UI same as ChatGPT's, it spits the whole answer at once instead of feeding it to you one word at a time, it has an alternative drafts button, the Google it button is a nice touch and it feels quicker.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#1141. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.
This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#115Earlier quoted context omitted.
What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?
The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#116Earlier quoted context omitted.
What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?
The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#117Earlier quoted context omitted.
Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.
And how are you going to distinguish those interactions from chatbots trying to sell you something?
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#1181. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#119By training model B using the results generated by model A, the copyright of corpus_A (OpenAI RLHF dataset) remains safeguarded, as model B is never directly exposed to corpus_A, preventing it from duplicating the content verbatim.
This process only transmits the concepts originating from corpus_A, which represents universal knowledge that cannot be claimed by any individual party.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#120Earlier quoted context omitted.
What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?
The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.
If the latter, there are many laws that say you can own an idea, provided it exists somewhere.