Live data from Hacker News

Wikimedia Enterprise – APIs for LLMs, AI Training, and More

enterprise.wikimedia.com

111–120 of 166 posts

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#111
post #28
post #15

Earlier quoted context omitted.

My understanding is that Enterprise is mostly about smoother delivery of content. E.g. Google has those informational summaries in search results, and it's easier for it to keep those current if there's a stream of updates it can subscribe to, rather than having to constantly download the full Wikipedia database dumps and parse them out. It also puts it all into conveniently formatted responses, and tries to do some…

> to keep those current if there's a stream of updates it can subscribe to, rather than having to constantly download the full Wikipedia database dumps and parse them out. I mean presumably prior to this they were using the [free] parsoid rest api along with the [free] event stream api. I highly doubt they were parsing the dumps. Its not even clear to me what the core value proposition of the new api is over the old…

There's also r&d going into additional, inferred data layers, e.g.: https://m.mediawiki.org/wiki/Wikimedia_Enterprise/Breaking_n...

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#112
post #17

Earlier quoted context omitted.

That's exactly what almost all knowledge is. When was it that you last verified something by yourself, with an experiment? You didn't test the things you know. You know things because you could see they were the consensus, and so you had no reason to challenge them. If an idea is disputed, then you trust it less. If it comes from a small number of reputable sources, then you trust it more than a large numbers of unre…

I thought knowledge, at least the best type comes from primary sources and from repeatable experiments with explicit premises as much as possible. This makes it sound like all knowledge is hearsay. If it is, what is the point of a place like Wikipedia or even an encyclopedia?

Indeed, the best type of knowledge comes from primary sources and original research. But those also produce an awful lot of not-knowledge.

Wikipedia's approach to sifting the knowledge from the not-knowledge is to prefer reliable secondary sources, i.e. sources deemed capable of telling the difference, mainly because they have a reputation for good editorial control. It's far from an ideal touchstone; but relying on "experts" is worse, because who's an expert? You need experts to identify experts, which is circular.

"Reliable secondary sources" doesn't amount to hearsay.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#113
post #75
post #74

Earlier quoted context omitted.

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineer…

It would be nice if we had a "lot of lawyers", given how frequently we're sued to try and get content censored, or having to fight orders to hand over user data - and more generally, how massive these new laws we need to comply with are (see, e.g., the EU Digital Services Act, which even creates an entirely new annual independent audit process).

We even intervene in other court cases to try and prevent bad laws being created/interpreted in ways that would hurt the open internet (see, e.g., our amicus in the French Constitutional Court two weeks ago, our lawsuit against the US NSA, and our amicus briefs in the two US "Netchoice" US Supreme Court cases). We also operate the https://foundation.wikimedia.org/wiki/Legal:Legal_Fees_Assis...

Sadly, we're a very tight team. The downsides of being a nonprofit...

Anyhow, I'm going to assume people are just ignorant as to how much WMF does, not deliberately trying to undermine it. https://meta.wikimedia.org/wiki/Assume_good_faith , as they say.

(disclosure: lawyer for WMF)

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#114
post #68
post #56

Earlier quoted context omitted.

Now imagine you were able to go and edit the sign that says it is ethanol free to add the details of your test and dispute the claim, that would improve the knowledge.

For that to make it into Wikipedia you'd have to first write an article in a reputable source.

Or rather, have a reputable source write an article about your work. If you write the news article about your own work and they publish it, the article is still a primary source (despite not being self-published)!

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#115
post #80

Earlier quoted context omitted.

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

1M$/year used to be enough. Source: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2...

Why stop there, the site cost $0/year back in 1999.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#116
post #90
post #80

Earlier quoted context omitted.

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

I don't agree that it's pure greed but hosting costs for the Wikimedia Foundation in the FY2022-2023 were $3.1 million. [1][2] [1] https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2... [2] https://wikimediafoundation.org/wp-content/uploads/2023/11/W... (page 4, pdf page 6 for expenses)

Hosting costs are £3m, but total expenditure is $160m - which obviously isn't covered by the interest on $250m.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#117
post #75

Earlier quoted context omitted.

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineer…

It would be nice if we had a "lot of lawyers", given how frequently we're sued to try and get content censored, or having to fight orders to hand over user data - and more generally, how massive these new laws we need to comply with are (see, e.g., the EU Digital Services Act, which even creates an entirely new annual independent audit process). We even intervene in other court cases to try and prevent bad laws being…

It isn't a question of the good work you do.

People care about Wikipedia, not the Wikimedia Foundation. The criticism arises from misleading advertising. WMF fundraising conflates the two, implying that _Wikipedia_ needs money or it'll die. Meanwhile the 2023 budget shows $3.1m in hosting expenses versus $24.4m in awards and grants.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#120
I'm reading the API docs https://enterprise.wikimedia.com/docs/

And they don't have an OpenAPI spec available to download? So they seriously expect developers to just manually write their own client code by reading and translating those example CURL commands by hand??!

Seriously it's 2024! Not having a spec to download and insepct for any API is a sign of incompetence. When tools like Postman or https://github.com/OpenAPITools/openapi-generator-cli exist and save hours of time, you can't seriously expect devs to write all this connecting code by hand anymore.

Post reply on HN