Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

11–20 of 319 posts

Re: Reddit is OpenAI’s moat

#11

Except Google and Microsoft/Bing search engines have already cached lots of that data just to provide their service. Facebook has its own treasure trove. Apple less so but they're also the least competitive in the AI space (of the FAANG) atm.

I'm no expert on corporate law, but maybe them aquiring reddit data makes that data illegal to use for others to train their models on? Even if available, if it's not legal to use the big guys can't touch it.

Re: Reddit is OpenAI’s moat

#12
I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?

Re: Reddit is OpenAI’s moat

#17

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

Seriously hope OpenAI is stripping all the canned meme replies that go on for hundreds of sub threads and end up as the top comments in threads before they train their models. How would you even do that reliably?

Re: Reddit is OpenAI’s moat

#19

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

You’re not alone in thinking that. The value in using Reddit as training data would be the question response format of threaded comments. The downside is that the vast majority of comments on Reddit are very low quality and repetitive. You’d have to do a lot of filtering to make it usable and what you’d be left with would be a much smaller pile of training data.

Re: Reddit is OpenAI’s moat

#20
I agree wholeheartedly with this ... and given Microsoft's trail of corporate bodies in its wake.. I wouldn't put it passed them to be orchestrating the cutting of all data-lakes for AI training, especially a clean and pre-processed source like Reddit. if data is water that corporations drink(which it is).. REDDIT is like finding a naturally occurring spring of Perrier water(by natural, I mean, we are the ants who bring this water one drop at a time from the soft petals of plants which collect the night dew).
Post reply on HN