Google+ cannot be used with customer or brand accounts anymore
1–10 of 62 posts
Re: Google+ cannot be used with customer or brand accounts anymore
#2Re: Google+ cannot be used with customer or brand accounts anymore
#3It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB.
Re: Google+ cannot be used with customer or brand accounts anymore
#4Re: Google+ cannot be used with customer or brand accounts anymore
#5The title should say: Google plus cannot be used.
Re: Google+ cannot be used with customer or brand accounts anymore
#6ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/
Re: Google+ cannot be used with customer or brand accounts anymore
#7ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/
arkiver and Fusl in particular have been absolutely amazing in what they've accomplished.
They also managed to pull in 94.5% of all Google+ Communities, which should provide the ability to view posts by Community (they're otherwise scattered among user posts). We're still assessing how much of that was login-page redirects in the last hour or two of the crawl, but it's amazing work.
I'd managed to send of about 80k larger, recently-active (100+ members, If you ever need to use that it's:
https://web.archive.org/save/
Where you replace "" with whatever it is you're trying to save, including the protocol string, say, this HN post: https://web.archive.org/save/https://news.ycombinator.com/item?id=19556665
That can be scripted, and my submissions used a bog-simple Bash script and xargs to plow through 100k submissions (20k appear to have been dead) in about 90 minutes, on very modest hardware.Also: the Internet Archive (and Archive Team) run off volunteers and donations. You can help, and please do.
https://www.archiveteam.org/index.php?title=Donate
(Not affiliated, but very grateful to them.)
Re: Google+ cannot be used with customer or brand accounts anymore
#8ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/
I'm really curious how such a large operation is legally pulled off, especially when services like Google+ intentionally try to make scraping difficult (and presumably for large offenders, will attempt a C&D).
Re: Google+ cannot be used with customer or brand accounts anymore
#9ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/
Re: Google+ cannot be used with customer or brand accounts anymore
#10ArchiveTeam did manage to scrape ~98.6% of the user profiles before it went down. It was around 16 hours from completion when Google pulled the plug. It was done with the distributed scraper 'Warrior' using a massive amount of small cloud instances to spread the load to 1000s of IP adresses. The dataset is ~1.45 PB. http://tracker.archiveteam.org/googleplus/
I really wonder what is so valuable about saving everything on the Internet. For all of human history until now, every little human interaction was fleeting and only meant something to the people involved. Things that were significant were preserved when the people involved decided they were important. Now we are trying to keep every one of those insignificant things, for what purpose? Training an AI is all I can thi…