Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

241–250 of 319 posts

Re: Reddit is OpenAI’s moat

#241
post #180

Earlier quoted context omitted.

My experience in the past 5-10 years on Reddit is: - the voting system is about what people want the truth to be, not the truth. - the users can be very easily gamed in comments to vote one way or another just by the initial voting being negative or positive (aka most just vote with the trend) - the opinions all come from a bias of urban and major metro area people. It's painfully obvious they don't understand any mi…

> the voting system is about what people want the truth to be, not the truth Probably the best example is legal type advice. Personal example: The whole situation with internet archive and their e-book lending process. I was vocal that they didn’t have a leg to stand on and would lose. The comment was just about the legal situation. I was downvoted to hell for that opinion and identified as some sort of book publishe…

I see the same thing with the situation surrounding the game "Dark and Darker". The TL;DR is that the game is likely stolen intellectual property and there is a lawsuit directed at the publisher and one of the developers. The lawsuit will likely prevent the game from being released, so the prevailing opinion on reddit is that the whole thing is an unjustified shakedown, despite the copious amount of evidence that the publisher stole intellectual property.

Re: Reddit is OpenAI’s moat

#243
post #236
post #107

I find the editorialization of my title hilarious. I did not put a ? at the end. The answer to any headline with a ? at the end is “no.” Whoever at HN edited it — this is not alright. Feel free to argue against the piece on its merits.

We sometimes put question marks on titles when a claim in a title is controversial. This is a longstanding practice: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... . > The answer to any headline with a ? at the end is “no.” There's a popular meme that says that, but it's not true. Edit: after reading everyone's objections I think the easiest thing to do is just reverse all of this. Sorry for the trou…

This sort of editorializing — especially considering that the article pertains to people that you are incentivized to please — is utterly unacceptable. At the least it should be clearly indicated as an edit by you.

In the strongest terms, this is unacceptable.

Re: Reddit is OpenAI’s moat

#244
post #166

Earlier quoted context omitted.

> I genuinely don't understand the appeal of AI for search If you're good at googling the flow is: Ask the question > Clock which result isnt spam and click it > Figure out how to dismiss the cookies gate without accepting the cookies > Dismiss the google login box > Dismiss the popover pushing you to install an app > Scroll the page or ctrl+F to find the answer With ChatGPT it's just type your question and your answ…

Aaand this is basically the story of silicon valley over and over again. Build a product that's more convenient, and people will use it. Google is so full of shit now, and even answer boxes are below 4 ads, it's just way more comfortable and efficient to ask chatgpt. This morning I wanted to know how many calories there are in a breaded chicken breast. Chatgpt told me in 3 seconds after asking. Google would have been…

out of curiosity, how did you verify the information?

Re: Reddit is OpenAI’s moat

#245

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

>You'd usually be better off training on trade publications, scientific journals, or fandom than Reddit. Bold claim. There are plenty of engineers and doctors on reddit answering nuanced questions that arent talked about in any journals. That is what is missing. The absurdly specific stuff that reddit gets. Sure you can answer with nonsense that sounds realistic, but you could also answer with the exact text that sol…

The GP ignored books for some reason.

But you get very different kinds of knowledge from each of those. They are not really comparable.

You won't find any useful explanation about how to solder surface electronic components in a PCB in a book, and you won't find a deep enough explanation about how to dimension an uncoupling capacitor for an amplifier in Reddit.

Re: Reddit is OpenAI’s moat

#246
post #243
post #236

Earlier quoted context omitted.

We sometimes put question marks on titles when a claim in a title is controversial. This is a longstanding practice: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... . > The answer to any headline with a ? at the end is “no.” There's a popular meme that says that, but it's not true. Edit: after reading everyone's objections I think the easiest thing to do is just reverse all of this. Sorry for the trou…

This sort of editorializing — especially considering that the article pertains to people that you are incentivized to please — is utterly unacceptable. At the least it should be clearly indicated as an edit by you. In the strongest terms, this is unacceptable.

I promise you that no one here is incentivized to please Reddit. Have you seen HN lately?

Same point if you mean OpenAI.

Titles on HN aren't the property of the submitter nor of the author of an article (though the author's opinion carries more weight). They're shared pointers and we edit them to be as neutral as possible. Mods do this all the time and have been doing it since HN began 16 years ago. It's primarily because of that practice that HN's front page looks completely different from Reddit's or any other similar forum's, and that's a crucial factor in the identity of the site.

If we didn't have this rule, people could (and would) use the title field to get away with anything. I'm not saying that's what you were up to, just that this is an important rule and bog-standard HN moderation.

Re: Reddit is OpenAI’s moat

#247

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

Given that sama is a board member of Reddit Inc, and that this is happening after GPT-4 was trained on Reddit data, I wouldn't jump to conclude they're upset at OpenAI. SO had publicly available, no-auth-required data dumps. This makes it difficult for them to know who is using their data. However, this surely isn't the case for Reddit who offered only API endpoints for this content, and I'm guessing you couldn't use…

While Reddit Inc didn’t provide data dumps, pushift.io published no-auth dumps of all site data back to 2008, which are still available on Academic Torrents.

Further, you could use the Reddit API to injest the full firehose of all site data in real time without violating rate limits. This is [one of the ways] how Pushift made their datasets.

Reddit was a lot more open that any other site approaching their size!

Re: Reddit is OpenAI’s moat

#248
I think the author is on to something here, but isn't there a bridge across the moat re: scraping?

Sure the API made things convenient, and scraping content will be a bit of an arms race, but scraping public postings that don't require a login to view still seems like a bridge over the moat. It's tempting to refer to the recent victory of LinkedIn over HiQ, but there's an important distinction in that ruling: It pertained to use of logged in accounts and did not explicitly overturn the prior ruling that public-facing data was fair game.

For now it's all still unknown territory. I would be skeptical of anyone who adamantly affirms a position that scraping for training is allowed/not-allowed because it is not yet settled law in the US where these companies are headquartered or the rest of the world where they operate.

Re: Reddit is OpenAI’s moat

#249
post #246
post #243

Earlier quoted context omitted.

This sort of editorializing — especially considering that the article pertains to people that you are incentivized to please — is utterly unacceptable. At the least it should be clearly indicated as an edit by you. In the strongest terms, this is unacceptable.

I promise you that no one here is incentivized to please Reddit. Have you seen HN lately? Same point if you mean OpenAI. Titles on HN aren't the property of the submitter nor of the author of an article (though the author's opinion carries more weight). They're shared pointers and we edit them to be as neutral as possible. Mods do this all the time and have been doing it since HN began 16 years ago. It's primarily be…

Look, I just don’t think you should editorialize without making it clear. Why is the claim controversial, for example?

Edit: I guess we went too deep in the thread, so we're resorting to edits. that was a rhetorical question. It's sort of like the Twitter community notes feature: making the edit, and its reason clear, builds trust. Otherwise, you just look like you're doing something shady (given that, again, the people in the article literally own HN).

In terms of YOUR opinion about the quality of my article – that is what upvotes are for. I guess you have a much heavier hand into what we're allowed to read than I realized! Again, given your incentives, that's worrying.

Re: Reddit is OpenAI’s moat

#250

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

>You'd usually be better off training on trade publications, scientific journals, or fandom than Reddit. Bold claim. There are plenty of engineers and doctors on reddit answering nuanced questions that arent talked about in any journals. That is what is missing. The absurdly specific stuff that reddit gets. Sure you can answer with nonsense that sounds realistic, but you could also answer with the exact text that sol…

Well, it depends on what you're asking the LLM to generate.

If I ask an LLM to write me an essay in the style of a caveman, the LLM simulates a redditor simulating a fictional caveman (who speaks english but with short words and bad grammar). After all, that's what was in the training data.

If I ask the same LLM to write me an essay in the style of a Harvard professor - who's to say the same thing doesn't happen?

So I can see how, when training a model, you might want it to know what's really written by a Harvard professor, and what merely professes to be.

Post reply on HN