Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

121–130 of 319 posts

Re: Reddit is OpenAI’s moat

#121
It is funny for me that you have to include statements like "I do not suggest conspiracy", or "I am not conspiracy theorists", just as if suspecting that corporations or governments are up to no good is a bad thing. Quite the opposite. It is common sense. The disinformation battle has been won on "their" end. Suspecting a "Conspiracy" is considered a bad thing.

Re: Reddit is OpenAI’s moat

#122
post #70
post #53

Earlier quoted context omitted.

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?

How does Google use Reddit's data in its models? You can access most (all?) Reddit pages without hitting Reddit at all via the "Cached" link in the search results.

Does Google have a special agreement with Reddit (and all other sites?) or is it legally "fair use" to reproduce web pages that are available freely online?

Re: Reddit is OpenAI’s moat

#123
post #69

Earlier quoted context omitted.

You are absolutely right minority or not. I actively avoid reddit because the majority of users there. The fact we're seeing government agencies start treading into using GPT models is frightening. We could find ourselves in a tragic comedy where all of the massive institutions and enterprises around us are addressing their serious issues via redditors by proxy.

depends of where you go on reddit. i've learned a lot of things for my hobbies. for example r/espresso and r/roasting are a source of good information. there are also places like r/askhistorians and many many otheres. reddit is not just r/funny.

niches and hobbies are dominated by beginners and ideas that can't be challenged (theres a word for this, I cant remember what it is). people aren't just talking about their experience on major subs though.

the opinions i run into real life can be very different then with people in the real world. the communities online are made up of the kinds of people who spend their time online, and the content you see on reddit is generally from the people who spend enough time on reddit that they want to browse new. these arent average people. only a minority professionals are actually engaged in reddit. even those ive seen run off because they don't agree with the acceptable opinions

Re: Reddit is OpenAI’s moat

#125
I love this theory from a conspiracy angle. But one issue is that they didn't cut off API access, they just made it more expensive. This creates a moat that small app developers can't cross, but not much of one for a funded AI company. If you that believed reddit's corpus was the key to training your own model, you'd just pay the money. It cuts off the wrong kind of person it would be targeted against if his theory is right.

Re: Reddit is OpenAI’s moat

#126
post #53

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

I think he's bullshitting even on API costs to maintain. They could just put it into reddit premium "want to use API ? Pay few bucks and use app of your choosing". Even $2 would easily cover the cost of lost ads and such. Then a much more expensive tier above for "data ingestors".

Re: Reddit is OpenAI’s moat

#127
post #53

Earlier quoted context omitted.

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

Yeah I think their real intention is to kill off third party Reddit apps so people are forced to use their own app, with all its tracking and garbage. Reddit on mobile browser is a case study of insane dark patterns Click to sort comments while not logged in? A popup appears asking to log in, with no close button. You have to click out of the box, but that’s not easily apparent View an 18+ subreddit? Let them browse…

This exactly, it’s unuseable and will only get worse. When “old” stops working it’ll be a sad, sad day. This blackout has made me realize how deeply I’ll miss reddit.

The thing is, to me, reddit would never be a viable enterprise without the volunteers moderation. If reddit had to pay for that they’d never ipo. And if they think the moderators will stay and be exploited they’re wrong. Thus if they go through with this they’ll fail as a public company. There is no there where they going.

Re: Reddit is OpenAI’s moat

#128
post #76
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

Yes. For anything from opinions of movies/games to DIY advice to nerdy stuff like the best watch to buy in a given price range or the best synthesiser. Or people's opinions on a particular episode of TV! Or basically any hobby. It's not great for business stuff because there are way more people using e.g. AWS for their hobby project than running multi-million $ businesses on it. What are the alternatives? Facebook an…

Agreed -- before making any major purchase I start at Reddit. The Internet as you say is a dumpster fire of blogspam top-10 lists where the top of the page is just STR."Updated in \{current_year}!"

Re: Reddit is OpenAI’s moat

#129
post #70
post #53

Earlier quoted context omitted.

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?

Reddit's terms are irrelevant. Unless Reddit requires a login to view its site (which would also prevent Google indexing), anyone can view the data without agreeing to the terms.

The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.

Re: Reddit is OpenAI’s moat

#130

The data Reddit has on its NSFW content (you know which one.. yes, the pornographic one) must be worth millions. No, no... jeez, you pervs... not the content. The upvote/downvote data for every image will probably train any generative AI on exactly what makes a good porn video/pic - and we all know sex sells. The dirty comments, and etc. On top of that, they'd make really useful leverage if you can connect any one of…

>The data Reddit has on its NSFW content (you know which one.. yes, the pornographic one) must be worth millions.

Blackmailing people who participated in /r/jailbait? I would bump that up a 0 or two.

Sorry, just thinking about how an AI would go about acquiring resources.

Please tell me that admin guy wasn't part of that.

Post reply on HN