Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

71–80 of 319 posts

Re: Reddit is OpenAI’s moat

#71

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

[deleted]

Re: Reddit is OpenAI’s moat

#72

I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?

Basically the only thing that comes to mind is that our collective views and knowledge about topics can change over time. If we had a hypothetical LLM in the 1930's, it would have had quite different views on things like mental health, civil rights, and so on.

But like you said, the models would ideally be trained on many other data sources as well.

Re: Reddit is OpenAI’s moat

#73
I feel like I’m taking crazy pills: Reddit is on CommonCrawl, which means its all available without API access, permanently from AWS’s CDN. These discussions are completely moot because the path of least resistance was already available and in use by OpenAI

Re: Reddit is OpenAI’s moat

#74

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

I suspect this is due to a fundamental misunderstanding of what tools like ChatGPT actually are, what their capabilities are, etc. I think there is a large population of people who think they're simply more sophisticated versions of Siri. In fact, I'd go so far as to say that the vast majority of people don't even realize they're powered by different things, but instead just see the "marketing term currently known as AI" as a singular monolithic entity, and all related technology just gets lumped into the same category. That's miles away from understanding how "use a language model to interpret what I'm asking you to search for, and then plug that into a normal search engine, and read me back the most relevant result", is fundamentally different from "use a language model to predict the stastically most likely next sequence of tokens based on the ones you gave me."

I don't mean this as a way of thumbing my nose at people who don't know this stuff, but rather, to point out what a massive failure the completely nonexistent consumer education surrounding ML products has been.

Re: Reddit is OpenAI’s moat

#75

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

This is the logical endgame for the search engine form factor. Remember watching mom ask google questions using complete sentences and punctuation with a “please and thank you”?

The appeal of a conversational response was too great to wait for a solution to the problem of training data quality.

Re: Reddit is OpenAI’s moat

#76
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

Yes. For anything from opinions of movies/games to DIY advice to nerdy stuff like the best watch to buy in a given price range or the best synthesiser. Or people's opinions on a particular episode of TV! Or basically any hobby.

It's not great for business stuff because there are way more people using e.g. AWS for their hobby project than running multi-million $ businesses on it.

What are the alternatives? Facebook and Discord are walled gardens. Quora is a shithole. The rest of the internet is blogspam and significantly more dubious than your average Reddit comment!

Re: Reddit is OpenAI’s moat

#77
post #33

This might as well be as good place as any to say that I wouldn't be surprised if in 20 years time we find out that the acquisition of GitHub by Microsoft was also done in OpenAI's favor. I know it's probably a crazy theory, but it did cross my mind a few months ago.

The really impressive feat here is not even that, but MS inventing time travel to acquire Github before OpenAI was a twinkle in Sam Altman’s eyes

Re: Reddit is OpenAI’s moat

#78
post #6

Could OpenAI acquire Reddit as an alternative to their IPO? Let's say it's worth $10B - quite a sum especially if OpenAI is 'only' worth $40B - $50B. But a world in which compute isn't a moat, LLMs aren't a moat... but real human-generated content IS a moat... maybe it could make sense?

I mean, that would be kind of insane, but it could actually work: OpenAI needs "clean", high-quality data as input for its AI; rather than harvesting data to sell to advertisers, they could and finance major internet infrastructure -- the equivalents of StackOverflow, Wikipedia, Reddit, Twitter, Facebook -- to harvest data to feed their models. That might actually make the world a better place.

Re: Reddit is OpenAI’s moat

#79
post #70
post #53

Earlier quoted context omitted.

I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…

> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?

It is OpenAI's and Microsoft's position that obtaining, using for training, and displaying data via AI is fair use.

Re: Reddit is OpenAI’s moat

#80

100% Open AI is already making fake AI-generated posts and comments in reddit for a while. Pack it up boys, internet ran by humans is over.

Should've packed it up back in usenet days: https://en.wikipedia.org/wiki/Mark_V._Shaney

Also, if reddit is in any dataset of these models then it's no wonder that I'm not intrigued by artificial """intelligence""" at all.

Post reply on HN