Live data from Hacker News

Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

modsandmembersblog.wordpress.com

371–380 of 424 posts

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#371
post #369

Earlier quoted context omitted.

I've had some preliminary discussions with EFF. There's a listing of online privacy and rights foundations I've compiled here: https://social.antefriguserat.de/index.php/Privacy_and_Elect... Many are mostly dead. Which may mean partly alive. (That Wiki in general might best be merged with AT's efforts. I'm the principle editor.)

I've passed on your idea. What did EFF have to say? Anything hopeful?

Basically, to write up and explain the idea, and it'll get passed on to the right people.

(At least one of whom is frequently seen on HN.)

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#372

Verizon claimed that the archivists violated the "terms of service" [1], but I couldn't find any reference to automation, downloading, crawling, or denial of service attacks that might apply. Does anyone have an idea of exactly what term or terms were violated by the archivists? [1] https://www.verizonmedia.com/policies/us/en/verizonmedia/ter...

Just playing a devil's advocate here. The way archivists are downloading the data can be said to disrupt the services, which is mentioned in the terms of service: 2. d. viii: "interfere with or disrupt the Services or servers, systems or networks connected to the Services in any way." I'd also like to point out that the apparent spokesperson Brenda Fowler said in her open letter to Verizon, that "If the problem is th…

Using the interface wouldn't block scrapers, yes? They do use the interface. But, this is academic I think. They offer a broken way to get our stuff, and say that we can't do anything else. Should we acquiesce to this?

As for bogging down the servers, my understanding was different from what the author said. They hadn't started to archive, but were in script testing mode and were accumulating yahoo accounts. What I saw of their activities, they were very careful about not overloading the servers. (I know that because I was backing up my own groups independently at the time, and I was able to do it. Luckily.)

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#374

This is why the change in ownership to private equity of .ORG TLD is problematic. What if the owners, or owners after it is sold further down the line, of the .ORG TLD prevent archiving? Or charge more for that? That would greatly affect web archive.org and Wikipedia.org. This move on preventing archival actions is probably setup to allow them to block it later after the sale goes through and say that it was a policy…

Wait. You're saying that the owners of the .org TLD could put restrictions on the use of those domain names? And something like "can't be used for archiving" is a legal possibility?

I admit to legal ignorance, but this does seem over the top.

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#375

Earlier quoted context omitted.

I meant for the sake of data protection, not for forensics. You start with all ones and gradually deplete your ability to write ones over time in electron charge memories such as SSDs.

This is not about companies following best practices but about what is going to happen when some of the supposedly deleted data pops up again, as it eventually will. Will a judge that is clueless about how computers really work consider that as a GDPR violation or not ? As deliberate or not ?

If you disable wear levelling, you could force it to end in a final all-zeros state.

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#376
post #98

Earlier quoted context omitted.

btw, maybe Mechanical Turk could help with the captcha part?

A couple of years ago I saw somebody giving a talk, where they demonstrated a CAPCHA-Solving API, with people from India solving the CAPCHAs for a few cents.

That's basically what the DeathByCaptcha server is.

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#377
If you have Verizon mobile service, call them!

In the US, dial 611, say "representative" at the prompt, then dial 3 for "something else."

I have Verizon prepaid, so switching is easy. I called and politely explained the situation to the customer service representative. I informed them that I will definitely switch away from Verizon if they delete the Yahoo Groups data without allowing archival. They promised to inform their manager and email me back.

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#378

Earlier quoted context omitted.

You can install ArchiveTeam Warrior: https://www.archiveteam.org/index.php?title=ArchiveTeam_Warr...

Thanks! I never heard of that before; just like project SETI though for archival purposes. What are the hardware requirements of that VM? I'm attempting to import it on my NAS4Free home NAS Virtualbox service which is the only machine I keep up 24/7 atm, but it takes forever to import. The hardware is very limited however (Atom D410 + a bit over 1GB RAM available), so I'm not sure it would succeed, but so far it load…

I’m running the Docker image on the smallest Hetzner VMs, with 5 concurrent groups and 40 shared rsync threads per container, and 12 containers per server. Start one container, do docker top on it to make sure it’s pulling, then start the others one by one, taking a few seconds between each to avoid overwhelming the CPU. I’ve got 6 of those little VMs going, and have rolled up 4GB and 2800 groups worth in 6 hours.

After they settle down, they’re more memory than processor intensive. I’ve considered playing with the settings a bit, but thought it was more important to get a bunch of them running on a couple different VMs at different sites.

If I were really feeling fancy, I’d write a nice deployment definition for orchestrating this with microk8s...

Re: Verizon/Yahoo Blocking Attempts to Archive Yahoo Groups – Deletion: Dec. 14

#380
post #237

Earlier quoted context omitted.

Who are the dark side of web scrapers?

People who personally have 100,000 Yahoo accounts because they made them back when you could just pretend to be blind and request the captcha in spoken form, and then fed it into Google's speech to text engine, fed it back in to Yahoo, made the accounts, and who also have a botnet of a million residential IPs and can spin up a bunch of servers to run some scrapers.

So spammers who had yahoo account mailers
Post reply on HN