Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

111–120 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#111
post #18

Thankfully archive.org exists, otherwise it would not be possible to get good training data in a few years when the internet is flooded with AI content.

Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM

Definitely not "most" of the internet. The internet is many exabytes at this point, while Common Crawl is only low petabytes.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#112
post #110

I don't care at all about this from a copyright or data ownership perspective, but I am a little skeptical that it's a good idea to be this incestuous with training data in the long run. It's one thing to do fine tuning or knowledge distillation for specialized domains or shrinking models. But if you're trying to train your own foundation model, is relying on output from other foundation models going to make them lea…

Where are any LLMs going to get data from as they become more ubiquitous and humans produce less publicly accessible original and thoughtful content?

The whole thing is a plateaued feedback loop.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#113

Earlier quoted context omitted.

Bard is only a week old and has a large "experimental" sticker on it. Besides its UI is better and the answers are succinct which I prefer.

They literally copied the Chatgpt UI, lol, only it looks like a dated Google UI. How do you prefer answers with less data?... that's crazy.

I just don't want to be hit with a wall of text every single time, it gets the point across with minimal padding (high signal to noise ratio), ChatGPT feels like it gets paid by the word and they do actually charge by token if you use the API.

As for the UI it's a take on the tried and true chat UI same as ChatGPT's, it spits the whole answer at once instead of feeding it to you one word at a time, it has an alternative drafts button, the Google it button is a nice touch and it feels quicker.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#114
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.

What shitty things are you talking about?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#115

Earlier quoted context omitted.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.

You can't own ideas, they got their own life-cycle.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#116

Earlier quoted context omitted.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.

I'm not necessarily arguing against you, but "problematic" is too generic a term to be useful. Genocide is "problematic". Having to run to the bathroom every 5 minutes to blow my runny nose is "problematic". What do you actually mean?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#117

Earlier quoted context omitted.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

And how are you going to distinguish those interactions from chatbots trying to sell you something?

A network of trust, backed by a social graph, which can be used to filter untrusted content.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#118
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

But remember many years back when it was news that Bing used Google search results to improve its results.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#119
This could actually be a good way to sidestep the training set copyright and access right issues. Copyright protection should solely encompass the expression of human generated content and not the underlying concepts.

By training model B using the results generated by model A, the copyright of corpus_A (OpenAI RLHF dataset) remains safeguarded, as model B is never directly exposed to corpus_A, preventing it from duplicating the content verbatim.

This process only transmits the concepts originating from corpus_A, which represents universal knowledge that cannot be claimed by any individual party.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#120

Earlier quoted context omitted.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.

Responding to sibling comment: We need some clarification here: are we speaking about just ideas in the abstract sense, or ideas that have been fleshed out i.e "materialized"

If the latter, there are many laws that say you can own an idea, provided it exists somewhere.

Post reply on HN