Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

51–60 of 319 posts

Re: Reddit is OpenAI’s moat

#51

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

It's funny that you mention recipes at the end because that's exactly what I like to use ChatGPT for. Every single recipe site has been so corrupted by SEO that you need to scroll past 12 paragraphs of nonsense to get to the actual recipe, and more often than not once you actually get to the recipe part you'll be bombarded with popups about newsletters and cookies or some late-loading ad will cause the view to shift past the step you're trying to look at.

ChatGPT is great for distilling that experience down to just a simple recipe for whatever it is I'm looking for.

Tangentially: props to AnyList for their amazing plugin that scrapes recipes off of sites like this and stores them in an easy-to-use format.

Re: Reddit is OpenAI’s moat

#52
people won’t care about Reddit’s internal governance policy when researching mattress toppers, but they will care about finding relevant, recent information. The gamble Reddit is taking is “do we think we can provide a competitive user experience and ecosystem without 3rd party developers?”. If the answer is yes, people will keep using and contributing to Reddit and it will remain a valuable source of information regardless of the lack of 3rd party apps. It could pay off, but if the Googles and Metas of the world just start offering AI training data as a product, it wouldn’t really benefit OpenAI if Reddit in particular locked down some older data at the expense of it’s user base.

Time will tell if the move will work out for Reddit in the long term.

Re: Reddit is OpenAI’s moat

#53

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched.

A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model.

You could get it down to a science where you only scrape any new data whenever you train the next model.

[0]: https://old.reddit.com/r/reddit/comments/145bram/addressing_...

Re: Reddit is OpenAI’s moat

#54
post #33

This might as well be as good place as any to say that I wouldn't be surprised if in 20 years time we find out that the acquisition of GitHub by Microsoft was also done in OpenAI's favor. I know it's probably a crazy theory, but it did cross my mind a few months ago.

Doesn't make any sense timeline wise, business execs at Microsoft were probably not thinking seriously about openai at the time

Re: Reddit is OpenAI’s moat

#55
This post is wrong on several aspects the other commenters have already pointed out, but it raises this tangential point which I'm curious about

"They could buy out more 3rd party clients, which would somewhat appease the community. This would be a terrible move IPO-wise."

Why would this be a terrible move IPO wise? If Reddit has two major apps, one for casual users and one for power-users, and it has enough control to eventually put (tasteful) ads in both, how is that bad for their IPO? I don't think it's likely, given how much goodwill they've already burned with the community, but from a pure IPO perspective it seems like a solid go forward strategy

Re: Reddit is OpenAI’s moat

#56

I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?

There are massive up to date archives available

https://tracker.archiveteam.org/reddit/

I guess either OP doesn't know about them or he assumes that AI researchers won't think of using them.

Re: Reddit is OpenAI’s moat

#57

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I don't think you're the minority. I want to actually read the thread and see what random individuals thought about a thing (with their different views intact), rather than getting an aggregate summary. Often the overlooked, 1 point comment (because he answered the question a month after the question was raised), is the one you're looking for.

Re: Reddit is OpenAI’s moat

#58

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

The appeal is that AI answers your question. Maybe a wrong answer. Maybe from questionable sources.

Search, on the other hand, doesn't answer your question.

Re: Reddit is OpenAI’s moat

#59
> They could extend the June 30th deadline, which I think is likely. This would get rid of 3rd party apps while, perhaps, keeping most of the community intact.

I don't get what the author is trying to say here.

If the deadline was extended, why would it "get rid of 3rd party apps"? I assume because some (Apollo, RIF) said they're going to close anyway even if the deadline was extended? If so, the community will not be "intact".

If the author think the death of 3rd party won't affect the community, then Reddit don't need to extend the deadline.

Re: Reddit is OpenAI’s moat

#60
post #30

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I think what the article tries to say is that OpenAI have already scraped Reddit for training data and with the recent API changes and subreddits going dark, new competitors in the AI space won't have it as easy to get the same training set.

Honestly this sounds like a shower-thought post. With even basic research, Internet Archive and The Eye have Reddit historical data freely available. My desktop PC has all comments and posts from 2007-early 2023, in a convenient jsonl zst. It's only 3TB.
Post reply on HN