ChatGPT and the Enshittening of Knowledge
61–70 of 301 posts
Re: ChatGPT and the Enshittening of Knowledge
#62Earlier quoted context omitted.
> It's true though, Wikipedia really is terrible and full of fake citations that lead nowhere. It's an anti-knowledge base that sometimes has good information. Yeah, Wikipedia is garbage puffed up beyond all belief. I literally just today saw something just like you describe. It should be viewed very skeptically on anything anyone disagrees over (because then it's just snapshots of an agenda-pushing battle).
Could you elaborate on this? If it is full of garbage a couple examples should be very easy to find. I completely agree that Wikipedia can have errors, but in topics that I am educated in it seems pretty decent and I can't remember the last time I came across any (comp sci for example). The most recent example I can think of is about is an article on vulture bees, and a citation about what their honey tastes like, wh…
I could give examples but I won't, because that would link my HN and Wikipedia accounts.
> So "garbage puffed up beyond all belief" and "full of terrible and fake citations that lead to nowhere" sounds a bit hyperbolic, tbh.
People unironically describe it as the "sum of all human knowledge," so it's definitely puffed up beyond belief. In reality, much of it is a slow battle of tendentious agenda-pushing, by people with weird personalities, played according to an arcane rule book (the first unstated rule of which is to never, ever acknowledge that you're pushing an agenda). That doesn't taint all of it, but it taints far more than you'd think.
Re: ChatGPT and the Enshittening of Knowledge
#63If you think of the knowledge base of the internet as a living thing, ChatGPT is a like a virus that now threatens its life. This is the same process SEO spam caused for search - it hampers the nature by which things function and the river needs to reroute (pagerank then usage metadata) to replace the lost signal. ChatGPT is more of an existential threat because it will propagate to infect other knowledge bases. Luke…
>And worse, then ChatGPT will digest its own excrement, worsening its own results further I wonder if we'll get a "dead sea effect" with AI, I've seen some stuff saying they've basically run out of high quality training data and now the training pool will get poisoned by AI generated shit. Basically garbage in, garbage out and these large language models might not be able to improve
Re: ChatGPT and the Enshittening of Knowledge
#64Legend.
Re: ChatGPT and the Enshittening of Knowledge
#65Re: ChatGPT and the Enshittening of Knowledge
#66Earlier quoted context omitted.
> All signs point to this strengthening the value of curation and authenticated sources. This is what they said about Wikipedia viz. Britannica… alas, it’s a brave new world out there… nowhere to run to nowhere to hide, see that Wiezenbaum post also on the homepage now, as another commenter quotes[0]: > Writing of the enthusiastic embrace of a fully computerized world, Weizenbaum grumbled, “These people see the techn…
It's true though, Wikipedia really is terrible and full of fake citations that lead nowhere. It's an anti-knowledge base that sometimes has good information.
Re: ChatGPT and the Enshittening of Knowledge
#67GPT is a language model, not an oracle. So my first take is that people querying it for research are doing it wrong. Then again, if there’s a large economic incentive to use it in that way, we are very well may end up with the kind of feedback loop that the author describes.
If it can't verify, it just won't answer/tickmark check the answers (happens 16% of the time... and ... always for maths). This is a feedback loop stopper, in the sense of only relying on your documents as the base, and being able to operate entirely without OpenAI (still though using other GPT models)
It's Fragen.co.uk - we believe that more answers formerly missed by CTRL+F will be found with this technology, than false answers taken as true. And if that's true, you are enbettering knowledge. And if not, you're enshittening it slower than the higher-hallucinating alternatives.
Re: ChatGPT and the Enshittening of Knowledge
#68If you think of the knowledge base of the internet as a living thing, ChatGPT is a like a virus that now threatens its life. This is the same process SEO spam caused for search - it hampers the nature by which things function and the river needs to reroute (pagerank then usage metadata) to replace the lost signal. ChatGPT is more of an existential threat because it will propagate to infect other knowledge bases. Luke…
If ChatGPT just consumes itself will the end result of all queries just eventually recurse down to a single answer? - for example "42"
[0] https://www.poeticous.com/shel-silverstein/hungry-mungry
Re: ChatGPT and the Enshittening of Knowledge
#69From my experiments, ChatGPT is at least valuable by generating boilerplate code for scripting languages. I believe once connected to enterprise codebase its quality and completion can dramatically increase.
> quality increase
We must have seen very different enterprise code.
Re: ChatGPT and the Enshittening of Knowledge
#70If you think of the knowledge base of the internet as a living thing, ChatGPT is a like a virus that now threatens its life. This is the same process SEO spam caused for search - it hampers the nature by which things function and the river needs to reroute (pagerank then usage metadata) to replace the lost signal. ChatGPT is more of an existential threat because it will propagate to infect other knowledge bases. Luke…
>And worse, then ChatGPT will digest its own excrement, worsening its own results further I wonder if we'll get a "dead sea effect" with AI, I've seen some stuff saying they've basically run out of high quality training data and now the training pool will get poisoned by AI generated shit. Basically garbage in, garbage out and these large language models might not be able to improve
Of course, some of those books will definitely be AI generated or garbage quality, and we all know many of those journal articles can be worth less than the paper they're printed on.
Yet even if we cut it down to 100,000 books and half a million scientific papers, that's a lot of training data each year... And that is just considering print media, there are other ways to get more content too.
For example, there is also transcription of video/podcasts/tv-shows/movies, etc. along with descriptions of the scenes for video, which could be used to generate a lot more stuff.
With people speaking to their devices and using text-to-speech more often, that's another source too--wouldn't be surprised if some devices just start recording conversations, and transcribing them.
Seems like a ton of potential data sources to me, although it will certainly get more difficult to cull AI generated stuff to prevent feedback, I'm sure the tooling will evolve to enable easy AI content detection and exclusion.