Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

151–160 of 319 posts

Re: Reddit is OpenAI’s moat

#152

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

[dead]

Re: Reddit is OpenAI’s moat

#153
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

That's really only an issue for technical topics. Stack Overflow would obviously be a much better source of quality training data for the kind of questions you presented. But for most things outside of that Reddit is by far the best online source that actually represents answers you would get from real people. Questions like what is the nightlife like in x city don't have a single true answer and thrive due to Reddit…

> Stack Overflow would obviously be a much better source of quality training data for the kind of questions you presented.

I dont think we need a llm that only responds with "Your question was marked as a duplicate"

Re: Reddit is OpenAI’s moat

#154
post #88

Earlier quoted context omitted.

Interesting question but sadly I am in no position to answer it. I think there are probably issues to address with scraping it blindly: - Can Reddit imprint its data somehow? A watermark? - Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization? If OpenAI can't work around this, I'm not sure they would be willing to cross any li…

I think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.

This sounds very dodgy. Will somebody be checking the degree of such "data alteration" and verify that the "AI" is actually inspired rather than copying?

To me this feels like its opening up the door for the elimination of copyright as any algorithmic layer interjected between scrapped data and end users could claim to be "inspired".

Re: Reddit is OpenAI’s moat

#155

If Reddit is such a valuable property it would make sense for one of the big AI players to come in and buy it out before it IPOs and the value shoots through the roof. Open Ai is the most likely one given the links between Altman and the management. But a new entry like maybe Apple could swoop in. We know how useless the management of Reddit are at making money from it, a sale in it's entirety might be the best exit…

Reddit is valuable for its milk, but gave away the cow years ago with the free API.

Re: Reddit is OpenAI’s moat

#156
Reading this makes me thanks the gods wikipedia is what it is (aka : a nonprofit, dedicated to fullfilling its original benevolent mission). Imagine if google or microsoft had succeeded in overthrowing wikipedia, what internet would look like.

Wikipedia, in itself, accounts for a good 50% of the positive image i have of the internet.

Re: Reddit is OpenAI’s moat

#157
post #80

100% Open AI is already making fake AI-generated posts and comments in reddit for a while. Pack it up boys, internet ran by humans is over.

Should've packed it up back in usenet days: https://en.wikipedia.org/wiki/Mark_V._Shaney Also, if reddit is in any dataset of these models then it's no wonder that I'm not intrigued by artificial """intelligence""" at all.

Looking at the examples discussing Reagan on the Wiki page, is this generated text looking to sway public opinion? This could be used to sway elections! What sort of precautions did the hacker named Mark V Shaney take before releasing this into the wild? These so called “Mark Off Chains” are too dangerous to be let into the wild. They should only be made available to the largest corporations who can protect us from their dangers. I’m going to write my senator to make sure this isn’t made publicly available!

Re: Reddit is OpenAI’s moat

#158

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

Eh I really strongly suspect this “you can use an LLM instead of google, I.e. as a knowledge model” to be a short lived trend. I hope Reddit sees it the same way. It’s kinda like using better bike infrastructure for sick cyberpunk roller derbies; a nice unexpected use case, sure, but its not built for that and sooner-or-later the issues will become all too apparent

IMO, the future will be using LLMs with live search results - which of course will probably require a funding model better than the terrible Display Ads setup we have for such content now. So best of luck to Reddit on that front - I hope your ipo fails

Re: Reddit is OpenAI’s moat

#159
post #88

Earlier quoted context omitted.

Interesting question but sadly I am in no position to answer it. I think there are probably issues to address with scraping it blindly: - Can Reddit imprint its data somehow? A watermark? - Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization? If OpenAI can't work around this, I'm not sure they would be willing to cross any li…

I think the bottom line is that Microsoft’s (and thus other for-profit AI initiatives) stance is that any and all data is fair game regardless of license or authorization. This results, in their opinion, from the fact that the AI alters the data, changes the output, and is otherwise “inspired” by the data in the same way an artist might be inspired by another without copyright infringement.

If my compiler was “inspired” by leaked Windows source code and altered it into a new form then I think their opinions on the matter would be very different.

Re: Reddit is OpenAI’s moat

#160

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

You are absolutely right minority or not. I actively avoid reddit because the majority of users there. The fact we're seeing government agencies start treading into using GPT models is frightening. We could find ourselves in a tragic comedy where all of the massive institutions and enterprises around us are addressing their serious issues via redditors by proxy.

>We could find ourselves in a tragic comedy where all of the massive institutions and enterprises around us are addressing their serious issues via redditors by proxy.

not with a bang, but with a "this. take my upvote my good sir"

horrific

Post reply on HN