Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

211–220 of 319 posts

Re: Reddit is OpenAI’s moat

#211
post #166

Earlier quoted context omitted.

> I genuinely don't understand the appeal of AI for search If you're good at googling the flow is: Ask the question > Clock which result isnt spam and click it > Figure out how to dismiss the cookies gate without accepting the cookies > Dismiss the google login box > Dismiss the popover pushing you to install an app > Scroll the page or ctrl+F to find the answer With ChatGPT it's just type your question and your answ…

Aaand this is basically the story of silicon valley over and over again. Build a product that's more convenient, and people will use it. Google is so full of shit now, and even answer boxes are below 4 ads, it's just way more comfortable and efficient to ask chatgpt. This morning I wanted to know how many calories there are in a breaded chicken breast. Chatgpt told me in 3 seconds after asking. Google would have been…

Also Google (and distressingly, DDG as of late) sanitizes the hell out of the SERPs to only present you with “approved viewpoints”.

And I’m not talking about fringe Q Anon type stuff: the other day I was looking up the specifics of China’s “Blue Sky Initiative” climate change policies and the only thing Goog/DDG would show me (despite several attempts at rephrasing the query) was Western industry think tanks bellyaching about how the policies effect profits. It took me a good ten minutes of refining my search before I got an English translation of the actual policy bullet points.

I can’t imagine this being such a PITA on 2013 era Google.

Re: Reddit is OpenAI’s moat

#212
post #64

> There is no question that Reddit is extremely valuable as training data. How often do you append “reddit” to your searches? Is it? When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations. Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic. For e…

> Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic.

That's kinda sorta exactly how ChatGPT answers stuff though, no?

Re: Reddit is OpenAI’s moat

#213
post #145

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

> Reddit is usually not the place to go for expert discourse Its not expert research but Reddit can be used to find probably-real-first-hand-experience on a given subject. Which is enough to prefix "reddit" to a lot of google searches. It just needs to be a marginal improvement to Google blogspam for it to have some intrinsic value.

I don't even think blogspam is the problem as much as SEO-gamed content.

Blogs are fine, I don't mind reading them.

The classic recipe that is 95% down the page with every possible word you could google in the preceding 95% in the form of a fake story.

Re: Reddit is OpenAI’s moat

#214

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

The author links this in the fourth paragraph; their argument is that while compute is not a viable moat, training data may be.

Re: Reddit is OpenAI’s moat

#216
post #107

I find the editorialization of my title hilarious. I did not put a ? at the end. The answer to any headline with a ? at the end is “no.” Whoever at HN edited it — this is not alright. Feel free to argue against the piece on its merits.

I really don't understand the policy of altering headlines here. It's such petty, obnoxious behavior! I really wish they would stop doing it and I would be really annoyed if it happened to an article that I wrote and submitted myself. Thank you for calling it out.

Re: Reddit is OpenAI’s moat

#217

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

One thing to consider is that while the large transformers used in LLMs might have these diminishing returns, we don't know what discrete jump in model architecture might come next. That model might gain a lot from even more training data. And might gain more from the semi structured data on reddit than the slightly less structured data on Wikipedia and Twitter

Maybe

Re: Reddit is OpenAI’s moat

#218
post #51

Earlier quoted context omitted.

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

It's funny that you mention recipes at the end because that's exactly what I like to use ChatGPT for. Every single recipe site has been so corrupted by SEO that you need to scroll past 12 paragraphs of nonsense to get to the actual recipe, and more often than not once you actually get to the recipe part you'll be bombarded with popups about newsletters and cookies or some late-loading ad will cause the view to shift…

I just used ChatGPT4 this past weekend to come up with smoothie recipes and then a shopping list to take to Whole Foods. Can also give it ingredients you already have and ask for recipes. It’s a killer application for me ha

Re: Reddit is OpenAI’s moat

#219
post #188

Earlier quoted context omitted.

- suddenly after the ChatGPT success, they realized that they have valuable data - next step is to stop third-party apps that generated these data - then they let the moderators show their power I’m not sure if people care about a CEO being exposed as a liar nowadays, but maybe some former Yahoo managers have another idea on how to destroy more value.

>former Yahoo managers Specifically the ones that ended up in charge of tumblr, so they can suggest "ban adult content on a platform famous for its adult content"

Part of reddit's API changes were going to prevent NSFW content from being viewed through the API but they reversed course on that.

Re: Reddit is OpenAI’s moat

#220
post #187

If you accept that Reddit could be OpenAI's moat, I think you could explain Reddit's behavior even without OpenAI's intervention. Like the premise here is that Reddit's data is super valuable and so OpenAI would want to stop others from getting it. But it also makes sense to say Reddit's data is super valuable so Reddit would want to limit access to it and be able to charge a premium for it. That said, I'm not totall…

> If you accept that Reddit could be OpenAI's moat, I think you could explain Reddit's behavior even without OpenAI's intervention. Like the premise here is that Reddit's data is super valuable and so OpenAI would want to stop others from getting it. But it also makes sense to say Reddit's data is super valuable so Reddit would want to limit access to it and be able to charge a premium for it.

I don't think I have a refutal, honestly. OpenAI aside, Reddit clearly has realized that their data is valuable, perhaps more valuable than their ads business. This is why my piece is frankly very speculative: the strategy aligns really well and the financial interests are there, but I wish I had a great argument why my theory is more likely than simply Reddit trying to capture the value of their data.

Post reply on HN