Live data from Hacker News

Wikimedia Enterprise – APIs for LLMs, AI Training, and More

enterprise.wikimedia.com

11–20 of 166 posts

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#11
post #5

> Real-time access to Knowledge Wikipedia knowledge is basically "democracy" knowledge, i.e. the more people decided to support an idea, the "truer" it gets. That's not knowledge at all!

That politicians get to scrub their pages shows there are cracks in places, but overall it's generally pretty ok

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#13
post #4

Looking forward to seeing the details on how they will handle revenue sharing with all of the people who contribute to them.

Funding that goes to editors (for example, stipends to travel to Wikipedia conferences) has been decreasing steadily each year. It's not due to a lack of money - Wikimedia consistently brings in millions of surplus revenue each year, see WP:CANCER[1] - so it's not clear that an additional revenue source will change anything.

^1: https://en.wikipedia.org/wiki/User:Guy_Macon/Wikipedia_has_C...

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#14
post #5

> Real-time access to Knowledge Wikipedia knowledge is basically "democracy" knowledge, i.e. the more people decided to support an idea, the "truer" it gets. That's not knowledge at all!

That politicians get to scrub their pages shows there are cracks in places, but overall it's generally pretty ok

> overall it's generally pretty ok

That's the kind of glowing praise I'd get from my 8th grade Geometry teacher when I got a C on a test.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#15

I was hoping more groups did stuff like this. The free sites doing it could handle some copyright issues if their EULA had a built-in license for distribution. In my previous analysis, (IIRC) I found that Wikipedia articles were under a copyleft license with attribution requirements. Does how Wikipedia Enterprise delivers this bypass that where neither use nor derivatives have those requirements? Or are they ignoring…

My understanding is that Enterprise is mostly about smoother delivery of content. E.g. Google has those informational summaries in search results, and it's easier for it to keep those current if there's a stream of updates it can subscribe to, rather than having to constantly download the full Wikipedia database dumps and parse them out. It also puts it all into conveniently formatted responses, and tries to do some signaling of content reliability.

The API responses do include information about the content license, which differs a bit between different wikimedia properties: https://helpcenter.enterprise.wikimedia.com/hc/en-us/article...

(I work for the WMF. I don't work on the Enterprise stuff, or have any insider knowledge of it. E.g. I used Google as an illustrative example, but I have no idea whether they're actually using this service. :D)

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#16
post #9

Yet this also exists: https://en.wikipedia.org/wiki/Wikipedia:Database_download

Which is a good thing. The entire corpus is CC-licensed and anyone can download it for free. If you want a real-time API, performance SLAs, machine parsable formats, support etc. then pay for it.

I don't have a problem with having paid special services, but the "machine parsable formats" is a bit troubling since I think that should be a core part of the open wikipedia project.

I submit this link after coming across this site while Googling for info on parsing wikipedia "infoboxes". I plan to check out their "Article Structured Contents (BETA)" API. Improving infoboxes to be machine-readable seems important. And it would be bad if didn't do this because it's a revenue stream for them.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#17
post #5

> Real-time access to Knowledge Wikipedia knowledge is basically "democracy" knowledge, i.e. the more people decided to support an idea, the "truer" it gets. That's not knowledge at all!

That's exactly what almost all knowledge is.

When was it that you last verified something by yourself, with an experiment?

You didn't test the things you know. You know things because you could see they were the consensus, and so you had no reason to challenge them.

If an idea is disputed, then you trust it less. If it comes from a small number of reputable sources, then you trust it more than a large numbers of unreliable people. So with the Wiki.

Human knowledge isn't from the platonic realm. Human knowledge isn't checked by a theorem prover. You get almost all of your knowledge from other people, and you have no choice but to trust them for almost all of it.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#18
post #5

> Real-time access to Knowledge Wikipedia knowledge is basically "democracy" knowledge, i.e. the more people decided to support an idea, the "truer" it gets. That's not knowledge at all!

Read up on Wikipedia's "reliable source" policies.

Information on Wikipedia is meant to be backed up by a verifiable source, partly to prevent a situation where knowledge only makes it onto Wikipedia if enough of the editors agree that it should be true.

Molly White made a great video and write-up explaining this a few months ago: https://blog.mollywhite.net/become-a-wikipedian-transcript/#...

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#19
post #14

Earlier quoted context omitted.

That politicians get to scrub their pages shows there are cracks in places, but overall it's generally pretty ok

> overall it's generally pretty ok That's the kind of glowing praise I'd get from my 8th grade Geometry teacher when I got a C on a test.

lol, I find wikipedia to be a bit all over the place or lack structure. It also is largely a giant block of text, though I do think the new LHS menu has improved things

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#20
post #4

Looking forward to seeing the details on how they will handle revenue sharing with all of the people who contribute to them.

The details are: none. There is no revenue sharing. That's literally the point of the Creative Commons license that Wikipedia uses. There's some things on the margins about attribution strings and share-alike requirements, but none of it would render AI training on Wikipedia illegal or compel AI companies to seek a separate, royalty-bearing license. Creative Commons is a "do what you wish" license, not a "free until I want money, then I rugpull you" license.

If you wanted revenue sharing, you were at the wrong party. You wanted the Microsoft-sponsored "Open Source is Communism" party down the block.

Post reply on HN