Live data from Hacker News

Google+ cannot be used with customer or brand accounts anymore

plus.google.com

1–10 of 62 posts

Re: Google+ cannot be used with customer or brand accounts anymore

#2
My Google+ notifications were more active than they have been in months as everyone was checking in to see how much longer was left. Google didn't specify the time of the shutdown, just "sometime on April 2nd", so none of us were sure when the cutoff would be. Middle of the day, as it turns out.

Re: Google+ cannot be used with customer or brand accounts anymore

#3
ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug.

It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB.

http://tracker.archiveteam.org/googleplus/

Re: Google+ cannot be used with customer or brand accounts anymore

#6
post #3

ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/

I really wonder what is so valuable about saving everything on the Internet. For all of human history until now, every little human interaction was fleeting and only meant something to the people involved. Things that were significant were preserved when the people involved decided they were important. Now we are trying to keep every one of those insignificant things, for what purpose? Training an AI is all I can think of.

Re: Google+ cannot be used with customer or brand accounts anymore

#7
post #3

ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/

Archive Team are absolute heros. The Google+ Mass Migration community learned of them and their "googleminus" project in January, and worked to help give information on the crawl, amount of data, and particulars of G+.

arkiver and Fusl in particular have been absolutely amazing in what they've accomplished.

They also managed to pull in 94.5% of all Google+ Communities, which should provide the ability to view posts by Community (they're otherwise scattered among user posts). We're still assessing how much of that was login-page redirects in the last hour or two of the crawl, but it's amazing work.

I'd managed to send of about 80k larger, recently-active (100+ members, If you ever need to use that it's:

    https://web.archive.org/save/
Where you replace "" with whatever it is you're trying to save, including the protocol string, say, this HN post:

    https://web.archive.org/save/https://news.ycombinator.com/item?id=19556665
That can be scripted, and my submissions used a bog-simple Bash script and xargs to plow through 100k submissions (20k appear to have been dead) in about 90 minutes, on very modest hardware.

Also: the Internet Archive (and Archive Team) run off volunteers and donations. You can help, and please do.

https://archive.org/donate/

https://www.archiveteam.org/index.php?title=Donate

(Not affiliated, but very grateful to them.)

Re: Google+ cannot be used with customer or brand accounts anymore

#8
post #3

ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/

Does the ArchiveTeam coordinate with services like Google+ for archival? Are the techniques used (such as spreading the load across 1000s of separate IPs) used for anything other than getting around Google's integrity services e.g. ratelimiting?

I'm really curious how such a large operation is legally pulled off, especially when services like Google+ intentionally try to make scraping difficult (and presumably for large offenders, will attempt a C&D).

Re: Google+ cannot be used with customer or brand accounts anymore

#9
post #3

ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/

I don't know if ArchiveTeam's bots can be logged in as Google Accounts, but I can still (for the time being at least!) browse around on a GSuite account, and an old grandfathered-in "Google Apps for Business" account

Re: Google+ cannot be used with customer or brand accounts anymore

#10
post #6
post #3

ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/

I really wonder what is so valuable about saving everything on the Internet. For all of human history until now, every little human interaction was fleeting and only meant something to the people involved. Things that were significant were preserved when the people involved decided they were important. Now we are trying to keep every one of those insignificant things, for what purpose? Training an AI is all I can thi…

I wonder too. I was berated on here for mentioning that I deleted my reddit comments when I quit reddit, as if I was somehow stealing (my own content) from the users. Just weird IMO.
Post reply on HN