Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

181–190 of 319 posts

Re: Reddit is OpenAI’s moat

#181
post #98

Earlier quoted context omitted.

Going outside is a great alternative to mindlessly consuming Reddit, but we're talking about researching opinions here. That's something that Reddit is actually good for.

Perhaps GP meant that since desire is the root of suffering, that rather than desire that your server work and suffer and strive to return it to functionality, you could go outside and mediate on the grass until you relinquished your desire for a working server. Or your desire for an umbrella that doesn't break in the wind, or the best public estimate of when Starship will launch, or any of the myriad other desires t…

Yeah, I could just meditate the entire day instead of doing anything. Why have desires at all?

Re: Reddit is OpenAI’s moat

#182

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

Why are they overlooking their biggest value as an organization is to sell limited truth to AIbros who want the _true_ up/down vote information, rather than the fake number they decide to provide to the plebians?

I think the problem is that the votes, not necessarily representative or anything “true” either.

Re: Reddit is OpenAI’s moat

#183
post #164

Data is the only moat in the AI age. The problem for businesses is that the output of models can be used to train other models, so even a massive data moat will eventually be eroded. Eventually all B2C AI companies will be just slick interfaces over open models that are too large to run on commodity hardware. B2B AI is in a bit better shape, both because compiling niche business data sets can be expensive, and there…

Current VRAM prices are like that, because of Nvidias greed. I assembled one of my previous gaming PCs 10 years ago and installing 32 GB of RAM wasn't a problem back then. But you can't even buy a consumer GPU with 32 GB of VRAM. Data center cards are considerably more expensive. There's no technical reason to not have 100+GB consumer GPUs today.

Historically consumer cards have been driven by the needs of gamers, and data center cards have mostly been slightly retooled consumer cards. Since game graphic needs have plateaued to some degree and I expect as AI gets incorporated into more things we'll see consumer cards that have lower game performance but a lot of memory and good basic model performance.

Re: Reddit is OpenAI’s moat

#185
post #70

Earlier quoted context omitted.

> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?

Reddit's terms are irrelevant. Unless Reddit requires a login to view its site (which would also prevent Google indexing), anyone can view the data without agreeing to the terms. The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.

Plus Reddit doesn't own the copyright to the posts, the users do.

Re: Reddit is OpenAI’s moat

#186
post #53

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

I'd also point out, it's not like the pricing of the API is going to be uniform for everyone anyway - spez has already said/conceded that certain users of the API can continue to use it for free (above the free limits). At that point it's a deliberate decision to make 3rd party apps pay the same price as AI companies, if they wanted to keep 3rd party apps around they could have set a different more reasonable price point for them or went another route (Ex. require 3rd party app users to have Premium).

Re: Reddit is OpenAI’s moat

#187
If you accept that Reddit could be OpenAI's moat, I think you could explain Reddit's behavior even without OpenAI's intervention. Like the premise here is that Reddit's data is super valuable and so OpenAI would want to stop others from getting it. But it also makes sense to say Reddit's data is super valuable so Reddit would want to limit access to it and be able to charge a premium for it.

That said, I'm not totally sold on the idea that Reddit has truly unique and valuable data that's a cut above what can be found elsewhere. After all, several of the so-called "glitch tokens" in GPT's tokenizer are Reddit user names that occur over and over in threads of people just counting. See: https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldm...

Re: Reddit is OpenAI’s moat

#188

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

- suddenly after the ChatGPT success, they realized that they have valuable data

- next step is to stop third-party apps that generated these data

- then they let the moderators show their power

I’m not sure if people care about a CEO being exposed as a liar nowadays, but maybe some former Yahoo managers have another idea on how to destroy more value.

Re: Reddit is OpenAI’s moat

#189
post #89
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

not only that but 80% of conversations on reddit is idiotic jokes and puns

This is absolutely true for any subreddit that shows up on the front page but many subreddits with smaller user bases or draconian moderation policies (like /r/askhistorians) can have excellent quality.

Re: Reddit is OpenAI’s moat

#190

Earlier quoted context omitted.

I think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.

If my compiler was “inspired” by leaked Windows source code and altered it into a new form then I think their opinions on the matter would be very different.

Not a great example: if Apple’s code leaked, theoretically they wouldn’t include it in the training as it’s not supposed to be seen by the public. If it’s public you can be inspired by it (so their logic goes).
Post reply on HN