Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

41–50 of 319 posts

Re: Reddit is OpenAI’s moat

#41

I agree wholeheartedly with this ... and given Microsoft's trail of corporate bodies in its wake.. I wouldn't put it passed them to be orchestrating the cutting of all data-lakes for AI training, especially a clean and pre-processed source like Reddit. if data is water that corporations drink(which it is).. REDDIT is like finding a naturally occurring spring of Perrier water(by natural, I mean, we are the ants who br…

This is much more YC/VC maneuvering than it is Microsoft.

Re: Reddit is OpenAI’s moat

#42
I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took:

- The entire corpus of data the community had curated over the last XX years

- The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data

- Potentially large amounts of traffic that would normally come to their sites via Google (e.g. site:reddit.com), that is now available instantly (and customized) via ChatGPT

Despite Reddit's probably closer connections to OpenAI than other startups through Y-Combinator and Sam Altman, I wonder how keen they are to actually work with a company that potentially destroyed a ton of their value, right before they were ready to IPO.

Re: Reddit is OpenAI’s moat

#43

I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?

> up to a certain date?

> just not recent enough

thats a very big deal for a lot of the content you find on reddit which users might want to get answers from from an AI.

Re: Reddit is OpenAI’s moat

#44

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

You are absolutely right minority or not. I actively avoid reddit because the majority of users there. The fact we're seeing government agencies start treading into using GPT models is frightening. We could find ourselves in a tragic comedy where all of the massive institutions and enterprises around us are addressing their serious issues via redditors by proxy.

[deleted]

Re: Reddit is OpenAI’s moat

#45

I agree wholeheartedly with this ... and given Microsoft's trail of corporate bodies in its wake.. I wouldn't put it passed them to be orchestrating the cutting of all data-lakes for AI training, especially a clean and pre-processed source like Reddit. if data is water that corporations drink(which it is).. REDDIT is like finding a naturally occurring spring of Perrier water(by natural, I mean, we are the ants who br…

I would like to be compensated for the digital oil I've produced.

Re: Reddit is OpenAI’s moat

#46

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

The appeal is that now regular search engines are so bad at giving you useful content that using a LLM is now the equivalent of "google-dorking" to find relevant information.

Re: Reddit is OpenAI’s moat

#47
post #6

Could OpenAI acquire Reddit as an alternative to their IPO? Let's say it's worth $10B - quite a sum especially if OpenAI is 'only' worth $40B - $50B. But a world in which compute isn't a moat, LLMs aren't a moat... but real human-generated content IS a moat... maybe it could make sense?

$10b is probably too low

Re: Reddit is OpenAI’s moat

#48
post #2

It would make a lot of sense for Microsoft to buy reddit: a curated source of data is the perfect moat! They could come as a white knight, restoring the API access but for apps only, conveniently blocking competitors while satisfying the users.

This is so smart that I'm now convinced it's what's happening.

Re: Reddit is OpenAI’s moat

#49
I’m not sure exactly what the moat would be here—the current data is already probably available, and future data OpenAI will have to pay for just like everyone else.

Ok, looking closer at the article I see

> The important piece is that it’s easiest for OpenAI to get the data (given that companies with co-investors help each other)

which seems like a pretty weak basis to form a moat (if Reddit IPOs as is their plan, sama’s influence will be not nearly as strong)

Re: Reddit is OpenAI’s moat

#50
(I know it's a satire)

It's quite naive if someone think all the major tech big boys who have at least moderate ambitions about AI haven't already at least archived Wikipedia, Reddit and StackOverflow.

Post reply on HN