Live data from Hacker News

Wikimedia Enterprise – APIs for LLMs, AI Training, and More

enterprise.wikimedia.com

71–80 of 166 posts

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#71
post #56

Earlier quoted context omitted.

When was it that you last verified something by yourself, with an experiment? Just a few weeks ago. I did a simple experiment to check whether the Super-94 at my local Chevron is indeed ethanol-free. It wasn't.

Now imagine you were able to go and edit the sign that says it is ethanol free to add the details of your test and dispute the claim, that would improve the knowledge.

This is nice analogy, but a wrong analogy. Wikipedia specifically does not allow original research.

https://en.wikipedia.org/wiki/Wikipedia:No_original_research

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#72
post #17
post #5

> Real-time access to Knowledge Wikipedia knowledge is basically "democracy" knowledge, i.e. the more people decided to support an idea, the "truer" it gets. That's not knowledge at all!

That's exactly what almost all knowledge is. When was it that you last verified something by yourself, with an experiment? You didn't test the things you know. You know things because you could see they were the consensus, and so you had no reason to challenge them. If an idea is disputed, then you trust it less. If it comes from a small number of reputable sources, then you trust it more than a large numbers of unre…

>That's exactly what almost all knowledge is.

That's what almost all human assumption and belief system is, also ideology and religion, but knowledge is indeed something different, and not of a type that should rely on democratic consensus. It instead needs to be held up by material evidence that's always subject to retesting no matter how unpopular a new idea is. This is obvious.

The rest of what you say could just as easily be applied to the foolish social dogmas of nearly any past age in human history, dogmas that so often turned out to be wrong. A small number of reputable sources (for their time) upheld doctrines such as geocentrism, religious extremism, hatred for certain racial groups and numerous fervent beliefs in the right of certain people to dominate others. These are just a few examples.

A more material one would be the certainty among reputable sources that plate tectonics were nonsense, until of course they were shown not to be by what started as an argument by only a few people who were deemed very unreliable.

None of this is to give weight to every crackpot idea put forth, or claim that all opinions are equally valid until stated otherwise, but what makes the difference is evidence, not consensus.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#73
post #66
post #44

Earlier quoted context omitted.

It shouldn't be much of a problem to ship the LLM along with attributions. (List of all sources used in the dataset - not a problem, unless they are secret, shady or illegal.) Wikipedia is one of the easy ones, since you need to attribute just 'Wikipedia' for the entire corpus instead of many individual users. The bigger issue is when a user uses the model to generate some text. Should they attribute it when using it…

Including a list of every Wikipedia author is possible but very inconvenient. https://en.m.wikipedia.org/wiki/Wikipedia:Reusing_Wikipedia_...

I wonder if it can be like those California proposition notices, just include it everywhere just in case

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#74
post #23

While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…

> While i am sympathetic to wmf finding alternative funding streams

Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat.

https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim...

$250m - they’re doing this out of greed, not need.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#75
post #74
post #23

While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineering of the sites on a technical level.

You also need HR, you need Finance, you need a lot of Lawyers, you need software developers, you need a travel department, a fundraising team, PR people, community relations people, grant-making for the extended open ecosystem around the Wikimedia movement, conference planning, and the list goes on.

You're off by enough to seem troll-ish at best.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#76
post #75
post #74

Earlier quoted context omitted.

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineer…

What are you talking about? The AVERAGE CD right now is 5%. My local CU is almost 6%. US bonds are currently ~4.5% - if you consider those unstable, I guess the US economy isn't stable - and if the US economy crashes, wikipedia will be the least of their or our worries.

Wikimedia's expenses are almost ENTIRELY going to staff. Their balance sheet for 2023 included $101m in expenses for salaries and benefits out of a total expense of $160m. Their hosting was $3m. So yes, I'm confident their network links and servers cost almost nothing, and they don't need anywhere near $101m in compensation to keep the lights on when the VAST majority of their content is contributed for free.

https://wikimediafoundation.org/wp-content/uploads/2023/11/W...

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#77
post #74
post #23

While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

Harvard is also exceptionally greedy to charge tuition, but also it wouldn’t be fair for legacy admits to get the brand and network without paying what is peanuts to most of those families.

LLM trainers need to pay for Wikipedia to help balance the information economy. If Google had to pay something substantial versus just lifting the content into their own “smart” results, then other sites would follow and wouldn’t have to rely on crappy SEO tricks.

Another perspective is that LLM trainers have largely been so disrespectful of IP / copyrights from their own greed that the content creators need to fight back with greed of their own. If the WM approach to media loses out to corporate and/or state control, it could e.g. make the Western internet much more like state-owned China.

Not entirely convincing arguments but it’s probably going too far to call WMF too greedy.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#78
post #76
post #75

Earlier quoted context omitted.

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineer…

What are you talking about? The AVERAGE CD right now is 5%. My local CU is almost 6%. US bonds are currently ~4.5% - if you consider those unstable, I guess the US economy isn't stable - and if the US economy crashes, wikipedia will be the least of their or our worries. Wikimedia's expenses are almost ENTIRELY going to staff. Their balance sheet for 2023 included $101m in expenses for salaries and benefits out of a t…

You may have missed them saying "when accounting for inflation". In the US at the moment that's around 3%. Thus your local credit union's savings account, a nice and stable investment, is effectively giving you around 3% appreciation in real-money each year right now. (I have no idea whether their broader point about the rate over-time is correct, admittedly.)

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#80
post #74
post #23

While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website?

Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

Post reply on HN