I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…
I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…
Reddit is OpenAI’s moat
71–80 of 319 posts
Re: Reddit is OpenAI’s moat
#72I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?
But like you said, the models would ideally be trained on many other data sources as well.
Re: Reddit is OpenAI’s moat
#73Re: Reddit is OpenAI’s moat
#74I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…
I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…
I don't mean this as a way of thumbing my nose at people who don't know this stuff, but rather, to point out what a massive failure the completely nonexistent consumer education surrounding ML products has been.
Re: Reddit is OpenAI’s moat
#75I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…
I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…
The appeal of a conversational response was too great to wait for a solution to the problem of training data quality.
Re: Reddit is OpenAI’s moat
#76> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…
It's not great for business stuff because there are way more people using e.g. AWS for their hobby project than running multi-million $ businesses on it.
What are the alternatives? Facebook and Discord are walled gardens. Quora is a shithole. The rest of the internet is blogspam and significantly more dubious than your average Reddit comment!
Re: Reddit is OpenAI’s moat
#77This might as well be as good place as any to say that I wouldn't be surprised if in 20 years time we find out that the acquisition of GitHub by Microsoft was also done in OpenAI's favor. I know it's probably a crazy theory, but it did cross my mind a few months ago.
Re: Reddit is OpenAI’s moat
#78Could OpenAI acquire Reddit as an alternative to their IPO? Let's say it's worth $10B - quite a sum especially if OpenAI is 'only' worth $40B - $50B. But a world in which compute isn't a moat, LLMs aren't a moat... but real human-generated content IS a moat... maybe it could make sense?
Re: Reddit is OpenAI’s moat
#79Earlier quoted context omitted.
I agree with you but not in entirety. This comment from spez[0] about blaming the API price changes on LLM's is too far fetched. A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model. You could get it down to a science where you only scrape any new data whenev…
> , that data can and will get scraped if the intention is to use for a model. How would that work from a legal perspective, though? Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
Re: Reddit is OpenAI’s moat
#80100% Open AI is already making fake AI-generated posts and comments in reddit for a while. Pack it up boys, internet ran by humans is over.
Also, if reddit is in any dataset of these models then it's no wonder that I'm not intrigued by artificial """intelligence""" at all.