Earlier quoted context omitted.
This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?
I don't agree that it's pure greed but hosting costs for the Wikimedia Foundation in the FY2022-2023 were $3.1 million. [1][2] [1] https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2... [2] https://wikimediafoundation.org/wp-content/uploads/2023/11/W... (page 4, pdf page 6 for expenses)
Wikimedia Enterprise – APIs for LLMs, AI Training, and More
101–110 of 166 posts
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#102Earlier quoted context omitted.
Wikimedia is Wikipedia's parent entity, but it's also the parent of: MediaWiki Wikibooks Wikidata Wikifunctions Wikimedia Commons Wikinews Wikiquote Wikisource Wikispecies Wikiversity Wikivoyage Wiktionary https://en.wikipedia.org/wiki/Wikimedia_Foundation
It sounds like an impressive list, but many of these are ghost towns and are of little value to machine learning. Wikiversity is a mess... Wikipedia is the crown jewel and probably the only thing of unique commercial value for ML.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#103While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…
There's a whole statement-of-principles thing that at least implies that the intention isn't to cannibalize the existing offerings: https://meta.wikimedia.org/wiki/Wikimedia_Enterprise/Princip... Though I imagine that only works so far as you feel you can trust the Foundation to stick to those principles, so that's complicated. :D There's also a bunch of FAQs here that sort of get at how the funding streams are suppo…
At the end of the day, WMF is made up of people, and people follow incentives. I'm not saying they are bad people, but they aren't saints either. They are just people like anyone else.
It might not happen today, but 5 or 10 years from now, i'm not so sure. Eventually there will be some situation where people involved will have to chose between something for the public good vs something that sells enterprise APIs better. If WMF becomes dependent on the enterprise money, it will be hard to chose the public good. When that day comes, the enshitification begins.
After all, google once claimed not to be evil. The motto didn't last.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#104Earlier quoted context omitted.
That's exactly what almost all knowledge is. When was it that you last verified something by yourself, with an experiment? You didn't test the things you know. You know things because you could see they were the consensus, and so you had no reason to challenge them. If an idea is disputed, then you trust it less. If it comes from a small number of reputable sources, then you trust it more than a large numbers of unre…
>That's exactly what almost all knowledge is. That's what almost all human assumption and belief system is, also ideology and religion, but knowledge is indeed something different, and not of a type that should rely on democratic consensus. It instead needs to be held up by material evidence that's always subject to retesting no matter how unpopular a new idea is. This is obvious. The rest of what you say could just…
Wikipedia doesn't establish consensus by pure numbers or voting, although it is a contributing factor. In disputes, it has moderated discussions with verdicts given by elevated users, including admins. Things like statistics and even (perhaps especially) precedent all weigh in. Popularity of a side can be weighted, but ruling purely based on popularity is actively discouraged.
This can lead to scenarios where 90% of users want something, but the moderater rules along with the 10%. Often, this happens when the discussion was initially among a bunch of relatively new users who aren't aware of some policy, and a more experienced editor points that a dispute is clearly not in line with some policy. This happens very regularly and is often a source of drama with long discussions.
This process actually arguably works better on popular and contentious pages; you get eyes and discussions of substance on those. Most boring pages are virtually ghost towns and are counterintuitively more susceptable to popularity-based consensus. Whatever you put up will likely stick, so it's just a matter of how many people and who will protect the page for the longest.
Also read this page: https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_vie....
The second page addresses your concern about not giving too much weight to fringe theories. It's not enforced as well as it could be in many places though; it can be hard to judge what's due or undue weight.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#105Earlier quoted context omitted.
It also seems Wikimedia isn't trying to relicense the content in any way that strips its e.g. CC-SA status, but rather providing the licenses as context alongside each API call. https://helpcenter.enterprise.wikimedia.com/hc/en-us/article... It's worth noting that https://creativecommons.org/faq/#artificial-intelligence-and... itself takes the general stance that "as a general matter text and data mining in the Unite…
> But as a practical matter, I wouldn't be surprised if some Wikipedia editors balk at their volunteer work being actively marketed and reformatted for ease of LLM training by the very platform that solicited their volunteer services, I think that will heavily depend on just what the money goes to. A better user experience, tightening up the code behind things, fewer nag screens for donations? Justifiable. Jimmy Wale…
Jimmy wales is not paid at all (he has a board seat, but that doesn't come with any money)
He of course has leveraged his fame from being "founder" quite extensively. I think most of his money comes from fandom which he is also one of the founders of.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#106Earlier quoted context omitted.
It also seems Wikimedia isn't trying to relicense the content in any way that strips its e.g. CC-SA status, but rather providing the licenses as context alongside each API call. https://helpcenter.enterprise.wikimedia.com/hc/en-us/article... It's worth noting that https://creativecommons.org/faq/#artificial-intelligence-and... itself takes the general stance that "as a general matter text and data mining in the Unite…
> I wouldn't be surprised if some Wikipedia editors balk at their volunteer work being actively marketed and reformatted for ease of LLM training As someone who avidly edited Wikipedia for 6-8 years, I am happy to see my volunteer work used for LLM training. I also agree some other editors likely aren't.
Redistribution of content is an entirely different matter, and the legal status of copyrighted material in relation to LLM training is an open issue that is currently the subject of litigation.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#107Earlier quoted context omitted.
Unfortunately this still results in plenty of hallucinations.
What do you mean? Have you or someone else already followed this exact approach?
I've attempted it as well a year ago (mostly for fun) for our project.
Yes, it can still hallucinate. But I would say it's much much much better in this regard than fine-tuning.
When I did it, the main issue was that our documentation wasn't exhaustive enough. There are plenty of things that are clear to our users (other teams in the company), but not at all clear to the LLM from the few text excerpts it receives. Also, our context was quite limited back then to just a few paragraphs of text.
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#108Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#109Earlier quoted context omitted.
> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…
This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?
Source: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2...
Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More
#110Interesting, the current DB dump is not for everyone, but if they also offer an LLM trained on Wikipedia data that answers and provide actual valida citation , please do. (not sure if duckduckgo stop offering that)