Live data from Hacker News

Wikimedia Enterprise – APIs for LLMs, AI Training, and More

enterprise.wikimedia.com

91–100 of 166 posts

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#91
post #76
post #75

Earlier quoted context omitted.

Over the long term of many years you're /lucky/ if a stable very-low-risk investment can net ~3% when accounting for inflation. Thus $250M could maybe net you roughly $7.5M/year. Exactly how many network links, servers, and engineering staff do you think that buys? It's way under what it operates on today, which is way under what it ideally should be for site like Wikipedia. And that's /just/ the operational engineer…

What are you talking about? The AVERAGE CD right now is 5%. My local CU is almost 6%. US bonds are currently ~4.5% - if you consider those unstable, I guess the US economy isn't stable - and if the US economy crashes, wikipedia will be the least of their or our worries. Wikimedia's expenses are almost ENTIRELY going to staff. Their balance sheet for 2023 included $101m in expenses for salaries and benefits out of a t…

An engineer costs $500k a year. Salary, benefits, office space, equipment, hr, legal, and other overhead. The engineer will only see a fraction of that, of course.

If you told me it took a hundred engineers to run Wikipedia I'd say, that's not totally unreasonable. Features, design, api, scaling, moderation, there's a ton for engineers to be doing.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#92
post #59
post #30

Is it’s ok to train AIs on CC-licensed data? What about attribution, should the AI be distributed with a note about the dataset’s origin? I could imagine a poorly trained model that returns content nearly identical to original; who would enforce CC rules then?

https://creativecommons.org/2023/08/18/understanding-cc-lice...

[deleted]

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#93
post #29

Earlier quoted context omitted.

CC-BY-SA to be specific. How does attribution of derivative works work in LLMs (etc)? When is the content they produce required to be CC-BY-SA as well (by virtue of the -SA part of the license)?

LLM output can't be copywritten as it's not the work of a human. The company would presumably have to include attribution to every piece of data the system is trained on for every query? Seems absurd, but then how would attribution work? Not really considered when the license was written.

Licenses go beyond copyright though, they tell you what you can and can't do with something someone else made. They're contracts.

I'm sure we'll be seeing more lawsuits (and likely new regulation and licenses) around this.

Are you saying LLMs are not derivatives of their source material? Why not?

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#94
post #80
post #74

Earlier quoted context omitted.

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

the expensive bit of serving big websites is the ads, the tracking, the analytics, the vast internal teams focused on endless avb testing etc. if you boil Wikipedia down it's mostly static pages with a crud editor. it's cached to the moon and back, most pages don't change. the pages are tiny. they're not paying aws for bandwidth.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#95
post #80
post #74

Earlier quoted context omitted.

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

Hosting 80GB of data? Absolutely. Even if it was 800GB.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#96
post #23

While i am sympathetic to wmf finding alternative funding streams, I do get nervous about these sort of things due to the inherent conflict of interest and incentives to canabalize the free offerings. I'm not saying that is happening now, but will it happen eventually? Additionally, originally it was promised this would all be open source, and officially they are sticking with that, but it seems like they are going w…

[deleted]

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#97
post #80

Earlier quoted context omitted.

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

the expensive bit of serving big websites is the ads, the tracking, the analytics, the vast internal teams focused on endless avb testing etc. if you boil Wikipedia down it's mostly static pages with a crud editor. it's cached to the moon and back, most pages don't change. the pages are tiny. they're not paying aws for bandwidth.

> most pages don't change

Generally the pages that are viewed a lot change a lot. Sure, there is a long tail of mostly static pages, but that is not super relavent.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#98
post #95
post #80

Earlier quoted context omitted.

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

Hosting 80GB of data? Absolutely. Even if it was 800GB.

Umm, its roughly 500 terabytes if you include uploaded files, but that is besides the point.

Hosting a bunch of static data is really easy but only a small part of running a site like wikipedia.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#99
post #93

Earlier quoted context omitted.

LLM output can't be copywritten as it's not the work of a human. The company would presumably have to include attribution to every piece of data the system is trained on for every query? Seems absurd, but then how would attribution work? Not really considered when the license was written.

Licenses go beyond copyright though, they tell you what you can and can't do with something someone else made. They're contracts. I'm sure we'll be seeing more lawsuits (and likely new regulation and licenses) around this. Are you saying LLMs are not derivatives of their source material? Why not?

I agree on lawsuits being likely and needed to establish precedent. CC-BY-SA specifies "other rights such as publicity, privacy, or moral rights may limit how you use the material".

At some point, maybe poorly-trained chatbots that consistently produce what's seen as avoidably/negligantly poor results may become regulated. Like how if a company poorly trains its employees, it is on the hook for their employees' behavior.

Re: Wikimedia Enterprise – APIs for LLMs, AI Training, and More

#100
post #80
post #74

Earlier quoted context omitted.

> While i am sympathetic to wmf finding alternative funding streams Why are you sympathetic to it? Their fund, at this point, can make enough off interest on a basic CD to not just pay for every possible cost they could have until the end of time, but make the maintainer(s) obscenely wealthy without breaking a sweat. https://upload.wikimedia.org/wikipedia/foundation/3/3e/Wikim... $250m - they’re doing this out of gre…

This is a rediculous thing to say. You think the interest off 250 million would be enough to run the wikipedia website? Do you have any experience operating web properties at scale to come to this conclusion? Did you base this on anything at all?

[deleted]
Post reply on HN