Live data from Hacker News

Large language models reduce public knowledge sharing on online Q&A platforms

academic.oup.com

41–50 of 366 posts

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#41
post #6

Well, if the users ask frequent/common questions to ChatGPT and get acceptable answers, is this even a problem? If the volume of duplicate questions decreases, there should be no bad influence on the training data, right?

They spoke to this point in the abstract. They observe a similar drop in less common and more advanced questions.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#42
post #5

It's a losing battle to try and maintain walled gardens for these corpuses of human-generated text that have become valuable to train LLMs. The horse has probably already bolted. I see this as a temporary problem however because LLMs are transitional. At some point it won't be necessary to train an LLM on the entirety of Reddit plus everything else ever written because there are obvious limits to statistical models l…

Only half true as maybe reasoning and actual understanding is not the strength of LLMs but it is fascinating that they actually can produce good info from everything they have read - unlike me who only read a fraction of that. Maybe dumb, but good memory.

So I think future AI has to read also everything if it is used like ChatGPT these days by average people to ask for advice about almost anything.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#43
post #30

Earlier quoted context omitted.

> At some point it won't be necessary to train an LLM on the entirety of Reddit plus everything else ever written because there are obvious limits to statistical models like this and, as a counter point, that's not how humans learn. You may have read hundres of books in your life, maybe even thousands. You haven't read a million. You don't need to. I agree but I think it may be privileging the human intelligence mech…

> These LLMs are polymaths that can spit out content at a super human rate. Do you mean in theory or currently? Because currently, LLMs make simple errors (eg [1]) and are more capable of spitting out, well, nonsense. I think it's safe to say we're a long way from LLMs producing anything creatively good. I'll put it this way: you won't be getting The Godfather from LLMs anytime soon but you can probably get an indust…

I think we’re in agreement. It’s going to take next generation architecture to address the flaws where the LLM often can’t even correct its mistake when it’s pointed out as with the strawberry example.

I still think transformers and LLMs will likely remain as some component within that next gen architecture vs something completely alien.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#44
post #8

Stackoverflow mods and power users being arseholes reduces the use of Stackoverflow. ChatGPT is just the first viable alternative.

It's an interesting question. The world has had 30 years to come up with a StackOverflow alternative with friendly mods. It hasn't. So the question is that has someone tried hard enough or can it be done it the first place. I am Stack overflow mod, dealing with other mods. There is definitely unnecessary hostility there, but most of question closes and downvotes Go 90% to low quality questiond which lack proper profe…

[deleted]

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#46
post #5

It's a losing battle to try and maintain walled gardens for these corpuses of human-generated text that have become valuable to train LLMs. The horse has probably already bolted. I see this as a temporary problem however because LLMs are transitional. At some point it won't be necessary to train an LLM on the entirety of Reddit plus everything else ever written because there are obvious limits to statistical models l…

> You may have read hundres of books in your life, maybe even thousands. You haven't read a million. You don't need to.

Sometimes online forums are the only place where you can find solutions for niche situations and edge cases. Tricks which would have been very difficult to figure out on your own. LLMs can train on the official documentation of tools l/libraries but they can't experiment and figure out solutions to weird problems which are unfortunately very common in tech industry. If people stop sharing such solutions with others, it might become a big problem.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#48
post #35

If a site aims to commoditize shared expertise, royalties should be paid. Why would anyone willingly reduce their earning power, let alone hand away the right for someone else to profit from selling their knowledge, unattributed no less. Best bet is to book publish, and require a license from anyone that wants to train on it.

Because it’s a marginal effect on your earning power and it’s a nice thing to do.

The management of these walled gardens will keep saying that to your face as they sell your contributions. Meanwhile your family gets nothing.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#49
post #8

Stackoverflow mods and power users being arseholes reduces the use of Stackoverflow. ChatGPT is just the first viable alternative.

It's an interesting question. The world has had 30 years to come up with a StackOverflow alternative with friendly mods. It hasn't. So the question is that has someone tried hard enough or can it be done it the first place. I am Stack overflow mod, dealing with other mods. There is definitely unnecessary hostility there, but most of question closes and downvotes Go 90% to low quality questiond which lack proper profe…

I think the problem isn't specific to SO. Text-based communication with strangers lacks two crucial emotional filters. Before speaking, a person anticipates the listener's reaction and adjusts what they say accordingly. After speaking, they pay attention to the listener's reaction to update their understanding for the future.

Without seeing faces, people just don't do this very well.

Post reply on HN