Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

231–240 of 319 posts

Re: Reddit is OpenAI’s moat

#231

I may be in the minority here, but if I want the opinion of Redditors on an issue, I will use a search engine to look for it specifically, thus knowing the provenance of the information I am receiving. I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all... Even though I am pretty sure it is already included in the training dataset already... I…

In my opinion, it comes back to search engines not giving you a the best answer. If you're looking for the best headset with whatever feature, you'll find mostly sponsored reviews, or at the very least hard to trust reviews.

On the other hand, some random people's opinions in a Reddit community with -apparently- no further agenda seem somewhat more honest.

Not that the answer is better but it gives you new data points in your search.

Basically, it's not one or the other, you can use both tools and that's probably why it makes sense to include Reddit in AI models (which do this job for you automatically)

Re: Reddit is OpenAI’s moat

#232
post #222

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

> you're ingesting expert knowledge that's highly specific and only discussed in a few forums. Reddit data could perhaps help with the latter This is exactly where I think Reddit's value lies. I disagree, though, that people don't go to Reddit for it. Here are a couple recent queries where I appended reddit to my search query: * best places to visit from London * best mattress sold in UK * should I bring king mattres…

> it's guaranteed (at least for now) that the answer came from humans who are not trying to sell you something.

No no no no! This is not guaranteed AT ALL on Reddit! Why do you think Redditors are constantly accusing each other of being shills? Because Reddit users are often corporate representatives doing "native marketing" or whatever they call it, and they are not upfront about it, because they are preying on the naivete of people like you!

Seeing the belief expressed that Redditors aren't trying to sell you something is.. distressing to me. Half of Reddit is some kind of ad, if not for a physical product it's for a political party or cause, or some celebrity's personal brand.

Please have a little cynicism online.. please, for the sake of everyone..

Re: Reddit is OpenAI’s moat

#233

Problem is, Reddit can’t obscure its content from Google because it is a source of traffic and new users for them.

Woah, this is a really good point. That brings up an interesting question: can Reddit stipulate in its TOS that a search engine can crawl the site for the sake of engine indexing, but not train its models on it? I don't see why they can't add such language.

Re: Reddit is OpenAI’s moat

#234

I feel like I’m taking crazy pills: Reddit is on CommonCrawl, which means its all available without API access, permanently from AWS’s CDN. These discussions are completely moot because the path of least resistance was already available and in use by OpenAI

Yeah, all the publicly available reddit data is freely available and the cat has been out of the bag for a while.

Reddit has a lot of other private information though. The entire corpus of moderated / censored comments is the internet's best alignment dataset. The real upvote/downvote numbers make for perfect RLHF.

Hell, Reddit could remove ads and replace them with paid RLHFing bots. Charge OpenAI money to selectively deploy AI bots to certain QA subreddits. All AI generated comments get deleted within a week. That way it doesn't lead to comment spam, but the bot still gets amazing upvote/downvote feedback and a ton of direct comment-reply feedback. 3rd party apps can't 'ad block' it either, since only Reddit knows which comments are fake.

(psst, reddit, if you're gonna steal my idea, I want to be paid for it)

Re: Reddit is OpenAI’s moat

#235
post #158

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

Eh I really strongly suspect this “you can use an LLM instead of google, I.e. as a knowledge model” to be a short lived trend. I hope Reddit sees it the same way. It’s kinda like using better bike infrastructure for sick cyberpunk roller derbies; a nice unexpected use case, sure, but its not built for that and sooner-or-later the issues will become all too apparent IMO, the future will be using LLMs with live search…

> the future will be using LLMs with live search results

N of 1, but I vastly prefer the Google Generative Results over chatGPT. Quality may not always be the same, but a “chat” seems like an awkward metaphor for finding info, and of course, chat GPT has no links to content when I’m worried about hallucinations.

I’m true google-search fashion, GenSearch has a lot less deep-in-the-weeds technical answers and will push you to simpler results. Eg. If you want to know what chemicals have a similar absorption properties to methane… you’re better off with ChatGPT or traditional search.

Re: Reddit is OpenAI’s moat

#236
post #107

I find the editorialization of my title hilarious. I did not put a ? at the end. The answer to any headline with a ? at the end is “no.” Whoever at HN edited it — this is not alright. Feel free to argue against the piece on its merits.

We sometimes put question marks on titles when a claim in a title is controversial. This is a longstanding practice: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que....

> The answer to any headline with a ? at the end is “no.”

There's a popular meme that says that, but it's not true.

Edit: after reading everyone's objections I think the easiest thing to do is just reverse all of this. Sorry for the trouble!

Re: Reddit is OpenAI’s moat

#238

Earlier quoted context omitted.

>most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. more than anything, AI/ML/DL enables people to make claims that cannot be refuted by any application of the scientific method. it's literally "not even wrong" in exactly the same way as people claim string theory is.

String theory wishes it had (unique) claims as easily falsifiable as "Once you have ingested several terabytes of natural language, you know how to generate natural language." Obviously you need a page of fine print to make a claim strictly falsifiable, but complaining about the absence of fine print in casual discussion is absurdly uncharitable unless you have reason to believe that agreeable fine print couldn't be…

>String theory wishes

are you a string theorist? are you sure that you understand the claims made by string theory and how falsifiable they are?

conversely

>fine print couldn't be drawn up and I'm 99% sure...

where does your 99% surety derive from?

As someone trained both in mathematical physics and DL, i'm here to burst your bubble: the falsifiability of both disciplines is exacty the same because they're both premised on knowledged derived of models with an astronomical number of free parameters.

just to be clear: i'm not poopooing LLN here, i.e. the models themselves are fine, i'm talking about absurd meta-claims (like ops about generating language) derived of observing many such models.

Re: Reddit is OpenAI’s moat

#239
post #107

I find the editorialization of my title hilarious. I did not put a ? at the end. The answer to any headline with a ? at the end is “no.” Whoever at HN edited it — this is not alright. Feel free to argue against the piece on its merits.

I really don't understand the policy of altering headlines here. It's such petty, obnoxious behavior! I really wish they would stop doing it and I would be really annoyed if it happened to an article that I wrote and submitted myself. Thank you for calling it out.

People get upset about title editing on HN because they only notice the cases they dislike. Any practice will seem obnoxious if you only count the things it gets wrong.

I can tell you after over a decade doing this job that the title moderation on HN is probably the single most beneficial thing we do. Without that, the front page would consist of linkbait, sensationalism, and rage. In other words, HN wouldn't exist.

Re: Reddit is OpenAI’s moat

#240

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

The only thing that didn't sit well with a lot of people about the leaked memo is that it ignored the quality of GPT4 vs GPT3 and made claims that all LLMs were poised to be on par, yet that isn't true till now.

What it also ignored (along with some of the comments here) is data ranking. Google didn't just build a search engine by crawling more of the web -- many search engines before it had already done that. Google managed to rank what's relevant and what isn't. Relevancy is hard. Similarly, not all scientific publications are ranked equally. Or for that matter, even publications with a lot of peer reviews or citations can become obsolete through new discoveries.

Reddit's data has value in that it can fill in a lot of the gaps left by more qualitative sources and furthermore the data is user-ranked by a trusted community. This also has implications for specialised querying, for example training on just r/fitness could be fairly useful for that community.

As a side note, other valuable data stores are not just text but voice/video as well. YouTube and podcast transcripts are readily available, for example to Google. Data and ranking is valuable all over again.

Post reply on HN