Live data from Hacker News

Large language models reduce public knowledge sharing on online Q&A platforms

academic.oup.com

261–270 of 366 posts

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#261

The problem is eventually what are LLMs going’s to draw from? They’re not creating new information, just regurgitating and combining existing info. That’s why they perform so poorly on code for which there aren’t many many publicly available samples, SO/reddit answers etc.

AI companies are already paying humans to produce new data to train on and will continue to do that. There's also additional modalities -- they've already added text, video, and audio, and there's probably more possible. Right now almost all the content being fed into these AIs is stuff that humans can sense and understand, but why does it have to limit itself to that? There's probably all kinds of data types it coul…

How would a custom language differ from what we have now?

If you mean obfuscation, then yeah, maybe that makes sense to fit more into the window. But it’s easy to unobfuscate, usually.

Otherwise, I‘m not sure what the goal of an LLM specific language could be. Because I don’t feel most languages have been made purely to accommodate humans anyway, but they balance a lot of factors, like being true to the metal (like C) or functional purity (Haskell) or fault tolerance (Erlang). I‘m not sure what „being for LLMs“ could look like.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#262
post #185

It’s been a relief to find a platform where I can ask questions without the fear of being humiliated Half joking, but I am pretty tired of SO pedantry.

I am very curious to see how this is going to impact STEM education. Such a big part of an engineer's education happens informally by asking peers, teachers, and strangers questions. Different groups are more or less likely to do that consistently (e.g. https://journals.asm.org/doi/10.1128/jmbe.00100-21 ), and it can impact their progress. I've learned most from publicly asking "dumb" questions.

I've found chatgpt quite helpful in understanding some things that I couldn't figure out when approaching pytorch for an internship

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#263
post #260

The problem is eventually what are LLMs going’s to draw from? They’re not creating new information, just regurgitating and combining existing info. That’s why they perform so poorly on code for which there aren’t many many publicly available samples, SO/reddit answers etc.

It may be an interesting side effect that people stop so gratuitously inventing random new software languages and frameowrks because the LLMs don't know about it. I know I'm already leaning towards tech that the LLM can work well with, simply because being able to ask the LLM to solve 90% of the problem outweighs any marginal advantage using a slightly better language or framework offers. Fro example, I dislike Pytho…

Alternatively, esoteric languages and frameworks will become even more lucrative ,simply because only the person who invented them and their hardcore following will understand half of it.

Obviously, not a given, but not unreasonable given what we have seen historically.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#264
post #39

Wondering about wider implications. If technical interactions reduce online, how about RL and how do we rate a human competence against an AI once society gets a habit from asking an AI first? Will we start to constantly question human advice or responses and what does that do to the human condition. I am active in a few specialized fields and already I have to defined my advice against poorly crafted prompt response…

> Will we start to constantly question human advice or responses and what does that do to the human condition. I'm surprised when people don't already engage in questioning like that. I've had to be doing it for decades at this point. Much of the worst advice and information I've ever received has come from expensive human so-called "professionals" and "experts" like doctors, accountants, lawyers, financial advisors,…

does this mean you trust complete randoms just as much?

if i need advice on repairing a weird unique metal piece on a 1959 corvette, im going to trust the advice of an expert in classic corvettes way before i trust the advice of my barber who knows nothing about cars but confidently tells me to check the tire pressure.

this “oh no, experts have be wrong before” we see so much is wild to me. in nuanced fields i’ll take the advice of experts any day of the week waaaaaay before i take the advice from someone who’s entire knowledge of topic comes from a couple twitter post and a couple of youtube’s but their rhetoric sounds confident. confidently wrong dipshits and sophists are one of the plagues of the modern internet.

in complex nuanced subjects are experts wrong sometimes? absofuckinlutely. in complex nuanced subjects are they correct more often than random “did-my-own-research-for-20-minutes-but-got-distracted-because-i-can’t-focus-for-more-than-3-paragraphs-but-i-sound-confident guy?” absofuckinlutely.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#266
post #17

If a site aims to commoditize shared expertise, royalties should be paid. Why would anyone willingly reduce their earning power, let alone hand away the right for someone else to profit from selling their knowledge, unattributed no less. Best bet is to book publish, and require a license from anyone that wants to train on it.

Why open source anything, let alone with permissive licensing, right?

This is a real problem with permissive licensing. Large corporations effectively brainwashed large swaths of developers into working for free. Not working for the commons for free, as in AGPL, but working for corporations for free.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#267
post #241

Eventually, large language models will be the end of open source. That's ok, just accept it. Large language models are used to aggregate and interpolate intellectual property. This is performed with no acknowledgement of authorship or lineage, with no attribution or citation. In effect, the intellectual property used to train such models becomes anonymous common property. The social rewards (e.g., credit, respect) th…

I don't understand this take. If LLMs will be the end of open source, then they will constitute that end for exactly the reason you write: > Large language models are used to aggregate and interpolate intellectual property. > This is performed with no acknowledgement of authorship or lineage, with no attribution or citation. > In effect, the intellectual property used to train such models becomes anonymous common pro…

I think that is because, overall, the human nature does not change that much.

You may be conflating several different media types and we don't even know what the lawsuit tea leaves will tell us about that kind of visual/audio IP. As far as code goes, I think most companies have already shown how they protect themselves from 'open' source code.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#268
post #35

Earlier quoted context omitted.

Because it’s a marginal effect on your earning power and it’s a nice thing to do.

"It's a nice thing to do" never seems to sway online platforms to treat their users better. This kind of asymmetry seems to only ever go one way.

Stack Overflow won't even let me delete my own content now that they're violating the license to it.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#269

Earlier quoted context omitted.

This being HN, I'd love to hear from one of the many IRC channel mods who literally typed (I'd guess copy/pasted) this kind of text into their chat room topics and auto-responders. If you're out there-- how does it feel to know that what you meant as a efficient course-correction for newcomers was instead a social shaming that cut so deep that the message you wrote is still burned verbatim into their memory after all…

> was instead a social shaming that cut so deep that the message you wrote is still burned verbatim into their memory after all these years? Maybe that was the point?

To be fair, after the ban expired, I started submitting the test cases as instructed and the community was very helpful under these constraints.

Re: Large language models reduce public knowledge sharing on online Q&A platforms

#270
post #72

Earlier quoted context omitted.

The first time that I asked a question on #cpp @Freenode was a unique experience for my younger self. My message contained greetings and the question in the same message. I was banned immediately and the response from the mods was: - do not greet; we don't have time for that bullshit - do not use natural language questions; submit a test case and we will understand what you mean through your code - do not abbreviate…

This being HN, I'd love to hear from one of the many IRC channel mods who literally typed (I'd guess copy/pasted) this kind of text into their chat room topics and auto-responders. If you're out there-- how does it feel to know that what you meant as a efficient course-correction for newcomers was instead a social shaming that cut so deep that the message you wrote is still burned verbatim into their memory after all…

> social shaming that cut so deep that the message you wrote is still burned verbatim into their memory after all these years

Oh my, this reminded me how some 20 years ago I was a high school kid and dared to install a more nerdy Linux distro (which I won't name here) on my home computer. After some big upgrade, the system stopped booting, and when in panic I asked for help at the official forum, I got responses that were shaming me for blindly copying commands from their official website without consulting some README files. That's how I switched to Debian and never looked back.

Post reply on HN