Earlier quoted context omitted.
Treating a human trafficker and someone who downloaded some files from a server the same is not in the job description of a prosecutor. What an absurd statement. It's very much the job of a prosecutor to make judgements about the severity of the crime and how to respond. And in this case, the prosecutor showed incredibly poor judgement. There wasn't even a particular reason why the case should go federal in the first…
The job of a prosecutor is to get a guilty verdict, the judge decides the sentence. At least that's how i understand it.
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
821–830 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#822Earlier quoted context omitted.
Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.
The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#823Earlier quoted context omitted.
Every scientific paper in the last 90 years or so is still under copyright, owned by the authors, the published, or the universities.
JSTOR was explicitly a library of public domain works, consolidated in a single place so that academic libraries could access those papers that nobody had an interest in distributing anymore. It recently added a bunch of copyrighted journals. It didn't have any of those at the time.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#824Earlier quoted context omitted.
I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.
Imo it’s not about you accessing things you want for free. If your family purchased a disc copy of the goonies before you were born and you watched it as a kid, your accessing of that content you wanted for free has no moral bearing. The core question is what impact does your consumption have, and I don’t think that participating in the streaming landscape is making things any better for anyone but their ceos.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#825Earlier quoted context omitted.
You actually don’t know that. The question would be, what proportion of human experiments are successful, and you don’t know the answer to that question, so the victims of experiments could dwarf the beneficiaries of successful research. That’s always the hard thing with basic utilitarianism.
1) It is an experiment. The point is to test something that nobody is sure about. Whether the result is expected or unexpected there isn't really such a thing as "unsuccessful"; even duplication of work is considered useful. 2) If you think the cost-benefits are bad, my advice is don't sign up to be experimented on. Nobody has to be experimented on if they don't want to be.
2) What if the families of experimentees receive payment? Allowing that would be a short way down the slippery slope from allowing the experimentation, and would make the matter of consent more difficult to assess.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#826Earlier quoted context omitted.
But the license doesn’t apply to me as a customer if I can’t be expected to even notice it. If I buy a book in a bookstore, no one would assume that training LLMs on it would be explicitly forbidden. And adding a note to the book would probably not be binding because no one is expected to read the legal notice in a book.
Ah, I assumed, that the clauses regarding the use in training of an LLM are printed inside the book somewhere.
There is nothing of value that the license gives me that I wouldn't already have if the contract didn't exist. I can already read the book, merely by having it in front of me.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#827Earlier quoted context omitted.
Swartz wasn't the kind of person to accept a plea bargain from an overzealous prosecutor who was indicting him on 13 felony charges with a possible sentence up to 35 years along with a $1 million fine. I assume he wanted a trial because he wanted to continue his fight for open access. And maybe he thought he might lose, but wouldn't lose on all counts, and would make the prosecution look unreasonable in the public ey…
His criminality is one matter, but the full weight of the Federal Government on him was an entirely separate matter. A federal prosecutor's job is to jail you regardless of whether it is for downloading a file from a server or for trafficking in humans, and they will come at you with the same vigor regardless of the crime. And nothing has changed about that.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#828a) Financed via inflation/"cantillon effect" due to ZRP/Stimulus that absolutely flooded the market with funny money in the hand of the sharks. b) Trained upon copyrighted work without compensation. c) Trained upon open source without even asking politely for authorization.
The Robber Barons from the last century can't even get close to our modern Feudal Tech Lords.
Unless you're one of us that have amassed multi-generation wealth in a exit in the last 20 years, you're completely fucked.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#829Earlier quoted context omitted.
Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.
The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.
Post-2008, ZIRP and QE pumped trillions into financial markets, making capital nearly free for those who could borrow at scale. That money didn’t go into raising wages; it went into inflating asset prices. If you owned stocks or real estate, you got richer. If you earned a paycheck, you watched housing and living costs go up while your wages stagnated.
VC was one of the biggest beneficiaries. With bonds yielding nothing, institutional investors had to chase returns, flooding venture funds with capital. That’s how we got an era of insane startup valuations, SoftBank-style mega-funds, and entire sectors built on free money. Growth-at-all-costs became the norm because the cost of capital was effectively zero.
Then COVID hit, and the Fed doubled down—more QE, more stimulus, even lower rates. Another massive wealth transfer. Money printer go brrr, asset prices moon, and suddenly we have SPACs, meme stocks, and a startup funding frenzy. Meanwhile, workers got a couple of stimulus checks, and by the time the dust settled, everything from rent to food to cars was way more expensive.
Now AI companies are running the same playbook that cloud megascalers ran before them—monetizing open-source work while locking out the people who actually built it. Cloud providers took open-source databases, infrastructure, and developer tools, turned them into managed services, and extracted billions in profit without meaningfully compensating the people who did the work. AI companies are now doing the same thing—scraping open-source repositories, academic papers, and public datasets, building models upon it then slapping on proprietary fine-tuning and charging for API access all the while blatantly raising capital by promising to make the same workers they stole from obsolte. All of it built on the backs of researchers, engineers, and artists who never see a dime, but also on the backs of everyone else via the cantillon effect.
Now rates go up, the bubble deflates, and who gets left holding the bag? Not the VCs who cashed out early. Not the bankers who took their fees. It’s the workers, the middle class, the open-source devs, and the late-stage startup employees who thought they had something real. The cycle repeats.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#830Earlier quoted context omitted.
What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.
> What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. Sure, agreed. > And the experience is equally mediocre. Absolutely not. I regret using a taxi nearly every time I opt for the cheaper option. It's only the "better" choice if you happen to be standing right in front of one. This experience is nearly universal no matter where I travel. I think people really forget how…