Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

101–110 of 319 posts

Re: Reddit is OpenAI’s moat

#101

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

ChatGPT like AI is the natural next step in the arms race of SEO vs Search. Right now is a mechanism to fight SEO. It will continue to work for a few years.

Re: Reddit is OpenAI’s moat

#102
post #6

Could OpenAI acquire Reddit as an alternative to their IPO? Let's say it's worth $10B - quite a sum especially if OpenAI is 'only' worth $40B - $50B. But a world in which compute isn't a moat, LLMs aren't a moat... but real human-generated content IS a moat... maybe it could make sense?

Why would you pay $10B for data which you can scrape for $50 from a few seedboxes?

Or if that's too much work for you, just download one of them many archives that are already freely available online.

Re: Reddit is OpenAI’s moat

#104
post #33

This might as well be as good place as any to say that I wouldn't be surprised if in 20 years time we find out that the acquisition of GitHub by Microsoft was also done in OpenAI's favor. I know it's probably a crazy theory, but it did cross my mind a few months ago.

The really impressive feat here is not even that, but MS inventing time travel to acquire Github before OpenAI was a twinkle in Sam Altman’s eyes

OpenAI certainly existed in 2018 but I don't think I noticed most people being aware of them outside of the industry.

Re: Reddit is OpenAI’s moat

#105
post #53

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

Yeah I think their real intention is to kill off third party Reddit apps so people are forced to use their own app, with all its tracking and garbage.

Reddit on mobile browser is a case study of insane dark patterns

Click to sort comments while not logged in? A popup appears asking to log in, with no close button. You have to click out of the box, but that’s not easily apparent

View an 18+ subreddit? Let them browse for 30 seconds, then ask they log in

Visit the site? Ask to login or use their app.

At this point, I’m more motivated to do anything but use their app if they’re this hellbent on getting me to download it.

Re: Reddit is OpenAI’s moat

#106
post #88
post #70

Earlier quoted context omitted.

> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?

Interesting question but sadly I am in no position to answer it. I think there are probably issues to address with scraping it blindly: - Can Reddit imprint its data somehow? A watermark? - Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization? If OpenAI can't work around this, I'm not sure they would be willing to cross any li…

I think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.

Re: Reddit is OpenAI’s moat

#107
I find the editorialization of my title hilarious. I did not put a ? at the end. The answer to any headline with a ? at the end is “no.” Whoever at HN edited it — this is not alright. Feel free to argue against the piece on its merits.

Re: Reddit is OpenAI’s moat

#108
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

That's really only an issue for technical topics. Stack Overflow would obviously be a much better source of quality training data for the kind of questions you presented.

But for most things outside of that Reddit is by far the best online source that actually represents answers you would get from real people. Questions like what is the nightlife like in x city don't have a single true answer and thrive due to Reddit's the diversity of thought.

Re: Reddit is OpenAI’s moat

#109
Data is the only moat in the AI age. The problem for businesses is that the output of models can be used to train other models, so even a massive data moat will eventually be eroded.

Eventually all B2C AI companies will be just slick interfaces over open models that are too large to run on commodity hardware. B2B AI is in a bit better shape, both because compiling niche business data sets can be expensive, and there won't be giant public data sets produced from their output.

Post reply on HN