Live data from Hacker News

Reddit is OpenAI’s moat

cyberdemon.org

271–280 of 319 posts

Re: Reddit is OpenAI’s moat

#271
post #217

The leaked Google memo "We have no moat, and neither does OpenAI" is instructive here: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language. The remaining misses are facts tha…

One thing to consider is that while the large transformers used in LLMs might have these diminishing returns, we don't know what discrete jump in model architecture might come next. That model might gain a lot from even more training data. And might gain more from the semi structured data on reddit than the slightly less structured data on Wikipedia and Twitter Maybe

I would argue that Wikipedia is much more structured than Reddit.

> we don't know what discrete jump in model architecture might come next

This is true, but as long as these models are based on creating sentences based on what sentences that it's seen have looked like, and not on fetching and understanding verifiable facts, these services will do more harm than good overall.

If anything, these models need two sources of training data:

1. The standard language model (as it is now) to be able to generate and process queries and provide understandable answers

2. A database of verifiable factual information that it can query in order to prevent it from completely hallucinating information and then asserting that it is verifiable factual information when asked [0].

Until we can solve the AI hallucination problem, these systems are going to require users to be much more careful with information they're given than most people can manage right now.

[0] https://www.cbsnews.com/news/lawyer-chatgpt-court-filing-avi...

Re: Reddit is OpenAI’s moat

#272

> Now, hear me out. This turn of phrase almost always is followed by an argument you don't have to bother reading because it is wrong.

Did you have any actual points from the article you wanted to discuss, or call out as wrong? Or are you just communicating that you refused to read it because of the first line?

Re: Reddit is OpenAI’s moat

#273

I think from Reddit's perspective, they are extremely upset with OpenAI, in the same way that I'm sure StackOverflow is upset -- OpenAI took: - The entire corpus of data the community had curated over the last XX years - The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data - Potentially large amounts of traffic that would normally come to thei…

Given that sama is a board member of Reddit Inc, and that this is happening after GPT-4 was trained on Reddit data, I wouldn't jump to conclude they're upset at OpenAI. SO had publicly available, no-auth-required data dumps. This makes it difficult for them to know who is using their data. However, this surely isn't the case for Reddit who offered only API endpoints for this content, and I'm guessing you couldn't use…

Is there anything with the new pricing approach that prevents Reddit from giving OpenAI more favorable negotiated rates that are not publicly disclosed?

Re: Reddit is OpenAI’s moat

#274
post #268
post #264

Earlier quoted context omitted.

> I guess we went too deep in the thread, so we're resorting to edits You can keep replying - you just need to click on the comment's timestamp to go to its page and the reply box will appear there. > given that, again, the people in the article literally own HN [...] Again, given your incentives, that's worrying. What people and what incentives are you talking about here? > that is what upvotes are for HN has never…

dang, just so I don't sound crazy: Michael Seibel is on the board of directors of Reddit. I don't know what sama's involvement is with YC anymore, but I would imagine that his wishes are taken pretty seriously in SV. This is why I think the editorializing should be made clear: even if your title edits are in good faith – which I totally buy – you want to protect community trust.

That's true, he is. (I forgot that in fact.) FWIW I've never discussed Reddit with him, and there's no pressure to moderate HN any particular way about these things. Our job at HN is to keep the community happy (er, as happy as possible) and to make HN as good as possible. Those are the things that make HN valuable to YC.

It should take only a brief glance at https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... to decide whether HN is protecting Reddit - the last two weeks have seen a tsunami of threads, almost all negative. Discussion of OpenAI isn't that lopsided but I wouldn't call it positive. sama isn't involved with YC in any way as far as I know. (Edit: I just asked, and it's a bit more complicated than that - he speaks at batch events occasionally, has ownership in past funds like most former employees do, invests in YC startups sometimes, etc.)

I appreciate the reference to good faith! that is kind of you. FWIW, the reason we don't mark titles as edited, or comments as edited when commenters edit them, or any of that kind of thing, is because HN's UI has always been minimal and in particular has always avoided ceremonial or bureaucratic trappings. This is the sort of quality that dies by a thousand cuts if you start adding details, and I don't think that would be the right tradeoff. It's just not in the overall spirit of the site. That doesn't mean we're trying to hide anything—we're always happy to answer questions about anything on HN, and I spend most of my time doing so.

Re: Reddit is OpenAI’s moat

#275
post #82

Earlier quoted context omitted.

I'm in full agreement, but I see an even better integration in other product lines (mostly gaming and search) Also, Microsoft doesn't have a generic social network to mine data from. They may prefer to stick to the professional world, but then acquiring a few high value sites just to close their API would be another move. How much $$$ do you think it would take for YC to say "hell yes!" to put HN under Microsoft cont…

> They may prefer to stick to the professional world, but then acquiring a few high value sites just to close their API would be another move. They already have LinkedIn, which is far more valuable than HN for this purpose.

Yes, Linkedin is a professional social network, but no, you will not find high quality content there.

There's a lot of value for keeping in touch with the zeitgeist (ex: sqlite innovations, llama etc) and finding new trends on HN

There's about 0 value from using linkedin.

Re: Reddit is OpenAI’s moat

#277
post #166

Earlier quoted context omitted.

Aaand this is basically the story of silicon valley over and over again. Build a product that's more convenient, and people will use it. Google is so full of shit now, and even answer boxes are below 4 ads, it's just way more comfortable and efficient to ask chatgpt. This morning I wanted to know how many calories there are in a breaded chicken breast. Chatgpt told me in 3 seconds after asking. Google would have been…

out of curiosity, how did you verify the information?

I didn’t, but it sounded right. I know it might be incorrect.

However, I’ve seen plenty of bad information in google answer boxes too. And finding it in actual search results is going to be way more time. It’s not a life and death question.

Re: Reddit is OpenAI’s moat

#278
post #276

Earlier quoted context omitted.

Andreessen Horowitz is probably the most famous VC firm in Silicon Valley…

Well, my downvotes are well deserved, asking such a question here.

If you were just asking a question you probably wouldn't have been downvoted though…

Re: Reddit is OpenAI’s moat

#279

Earlier quoted context omitted.

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

> I genuinely don't understand the appeal of AI for search If you're good at googling the flow is: Ask the question > Clock which result isnt spam and click it > Figure out how to dismiss the cookies gate without accepting the cookies > Dismiss the google login box > Dismiss the popover pushing you to install an app > Scroll the page or ctrl+F to find the answer With ChatGPT it's just type your question and your answ…

I mean the internet lies and makes stuff up too, being first on Google results has nothing to do with truthfulness.

Re: Reddit is OpenAI’s moat

#280
post #114

Earlier quoted context omitted.

I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. I almost never accept a single search result as authoritative unless I'm pretty familiar with the source. Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in…

> I genuinely don't understand the appeal of AI for search. The provenance of information is just as important as the resulting information for pretty much anything I search for. Search requires work on the part of the user to distinguish between good links and bad. AI is an oracle just tells you what you're looking for. Now you and I might think this is a terrible way to evaluate the veracity of information. But thi…

> AI is an oracle just tells you what you're looking for.

LLMs are oracles that arrange words in a probabilistic order that are grammatically correct and may be factually correct. Unfortunately there's no way to evaluate the probability of confabulation with any of the LLM chat bots. The distribution of occurrences confabulation is also not regular or predictable nor is it fixed. So you can't ever say "ChatGPT is bad at X" because it can be bad today, good tomorrow, then bad the next day.

Post reply on HN