Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

671–680 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#671

Earlier quoted context omitted.

> If you reject Macaulay on copyright because he was an imperialist On the contrary I would argue that this is precisely why you SHOULD NOT take his opinion on copyright. One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. I did not quo…

if Indians are so free from colonialism, why are their parents forcing them to choose between medicine or tech, simply so they can get a job on the antipode of where they are born??

Because the wealth was transferred from India to the "antipode" through Colonization. GDP reduced from 25% Pre-Colonization (and 30% if you take Pre-Islamic Colonization) to merely 4% Post-India's Independence. At least Indians are not reverse-colonizing the West.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#672
post #590

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

And if you were in any doubt before, this lesson is now exemplified by the holder of the highest office in the land and approved by popular vote. The rewards of acting ethically are, unfortunately, sometimes only personal. This must be a hard environment to raise children in, given the examples they see around them.

Parent here: it takes a lot of discussion, but it's a great time to talk about the reality of evil and villains. My kids are on the good side, or at least I like to think so ...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#673

Earlier quoted context omitted.

Taxi Medallion laws were also a Reputation Engine that was publicly queryable , subject to FOIA laws and generally had easy to search public databases for them, with detailed notes. Sure Uber/Lyft boil that into a "friendly" 5-star UI, but do you have any idea what data contributed to that star rating? Do you always trust the algorithms that compute them from a bucket of metrics you can't directly request? Sure, Meda…

> Taxi Medallion laws were also a Reputation Engine that was publicly queryable , subject to FOIA laws and generally had easy to search public databases for them, with detailed notes. You just landed at the airport and need a cab. You fax your FOIA requests for each of the hundred cab companies in the area, which they're required to provide within 20 business days. Your return flight is in 3 days and it would be nice…

> You just landed at the airport and need a cab. You fax your FOIA requests for each of the hundred cab companies in the area, which they're required to provide within 20 business days. Your return flight is in 3 days and it would be nice to leave the airport before then.

If you just landed at the airport, you rely on police enforcement keeping bad actors from having medallions. The medallion itself is the primary "this person is a reputable cab driver". That's also entirely why the Regulatory Capture in some cities was so effective in controlling supply of medallions, because it was city police enforced.

Many cities required taxis to have their medallion number painted on the outside, and there were phone numbers you could quickly call (in the days of payphones even) to get quick information about a medallion or to report a complaint/problem with one.

Today a few cities have updated that external paint requirement (and inside the car medallion papers) to include QR codes for even quicker lookup on modern phones or to even use an app to do nice things like pay for the Taxi without needing to broker/negotiate it. Those kind of technological improvements have kind of gotten lost in the wash of the speed of which Uber/Lyft moved fast and broke things, but were always possible.

> So compete with them instead of banning them. Fund an open source ride hailing app with open data. Don't require anyone to use it. If it's better, they will. If it's not better, why should they be forced to?

The history of taxi companies say that they are only as open as they are forced to be. I never said anything about banning Uber/Lyft. Competition is not the problem; destroying public safety regulations in the name of competition is the problem. I said that Uber/Lyft should have been required to do the same or similar paperwork that medallions represent, that both of their data should be open under the previously existing laws, as a public good. Break the artificial scarcity, sure, give Uber/Lyft a license to "print medallions" if that breaks existing Trusts. But get that data open and available to the public (and enforceable by the public's law enforcement). Neither would want to do that because their rating systems are secret sauce and "competitive advantage", they would need to be coerced by regulations. That's what regulations are for, the public good that competition doesn't care about/can't care about/needs to keep "secret sauce" for advantages.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#674
This reminds me of Peter Sunde's "komimashin"

https://www.engadget.com/2015-12-21-peter-sunde-kopimashin.h...

It's obviously absurd to enforce copyright as bytes are copied around instead of as it is used. Training an LLM is a different thing than re-hosting and giving away copies to other people.

If you don't want people to transform your works - keep them private. You don't own ideas.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#675

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I sometimes think my adblocker should very much lie to the page that "yeah, watched that, totally" in an undetectable way.

https://adnauseam.io/

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#676

Earlier quoted context omitted.

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Sure creative people will always create but the scope of that creativity will be limited if we do away with intellectual property. Steve Spielberg would probably always have created movies, but he wouldn't have been able to make Jurassic Park, Saving Private Ryan,or Indiana Jones without capital from the studio system, and the studio system wouldn't have provided him with that capital of they couldn't extract economi…

They could have started a crowdfunded project and might still have made a great movie. If people truly like the created movie, why noch fund another one? Only no one would be paid millions for acting most likely, and that would be fine.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#677

I wonder what happened to the related OpenAI training GPT3 on the books3 dataset story[1] from ~2 years ago? [1]: https://www.wired.com/story/battle-over-books3/

I think this one is different because the legality of training on copyrighted material is an open legal question while distributing/seeding copyrighted material is decidedly illegal.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#678

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

No post body was provided.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#679
post #590

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

And if you were in any doubt before, this lesson is now exemplified by the holder of the highest office in the land and approved by popular vote. The rewards of acting ethically are, unfortunately, sometimes only personal. This must be a hard environment to raise children in, given the examples they see around them.

This argument may focus too much on the category of external rewards.

I might well be kidding myself or self-justifying, but I believe internal rewards are at least as important. Some materially successful people are deeply unhappy.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#680

Earlier quoted context omitted.

It's a library of historical scientific work. You will find the famous Einstein's 3 1905 papers there, for example.

Every scientific paper in the last 90 years or so is still under copyright, owned by the authors, the published, or the universities.

JSTOR was explicitly a library of public domain works, consolidated in a single place so that academic libraries could access those papers that nobody had an interest in distributing anymore.

It recently added a bunch of copyrighted journals. It didn't have any of those at the time.

Post reply on HN