Live data from Hacker News

Getting 10TB of GitHub logs and extracting details of all users and repositories

trickest.com

31–40 of 61 posts

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#31
post #17
post #8

Earlier quoted context omitted.

I have a GitHub account under my real name, but recently I've started using GitHub under a couple of other names instead. There's so much stuff you do in public on GitHub that I want to avoid people doing exactly this kind of analysis on. I wish using multiple identities was at least some level of foolproof though. I have to be careful to configure my local copies of repos to use the correct username, masked email, a…

note that this is technically against their TOS if not using paid accounts: > One person or legal entity may maintain no more than one free Account (if you choose to control a machine account as well, that's fine, but it can only be used for running a machine) https://docs.github.com/en/site-policy/github-terms/github-t...

They'll have to catch me first.

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#32

I'm a big fan of neat statistical analyses, but when I look at the archive site https://www.gharchive.org/ , the overwhelming feeling I have is "creeped out". Taking periodic snapshots of repositories and their issues and wikis sounds good, but do we really need a log of every time someone watches an issue, and every commit message being irrevocably set on public record? That level of details on individual activity s…

Don't blame the archiver who is only doing what is allowed of them by Github. If you don't want to be public don't be public. Other archivers aren't making themselves publicly known but have the data all the same.

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#34
post #17

Earlier quoted context omitted.

note that this is technically against their TOS if not using paid accounts: > One person or legal entity may maintain no more than one free Account (if you choose to control a machine account as well, that's fine, but it can only be used for running a machine) https://docs.github.com/en/site-policy/github-terms/github-t...

internet would be better with total anonymity.

That's one of my favorite parts of reddit. Accounts are just pseudonyms, and you can generate as many as you want. I personally generate a new one every few months, which helps keep too much identifying data from building up over time.

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#36

I'm a big fan of neat statistical analyses, but when I look at the archive site https://www.gharchive.org/ , the overwhelming feeling I have is "creeped out". Taking periodic snapshots of repositories and their issues and wikis sounds good, but do we really need a log of every time someone watches an issue, and every commit message being irrevocably set on public record? That level of details on individual activity s…

Lol let me introduce you to a little organization called the National Security Agency, with their "creepy" periodic snapshots of much more intriguing datasets. "Stellar Wind" is a good place to start. Including, but of course not limited to, every communication made by any person within the United States (or outgoing) for the better part of two decades. Internet traffic, communications, all of it. https://oig.justice…

Looks like I'm destined for the gulag for sure.

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#37

I'm a big fan of neat statistical analyses, but when I look at the archive site https://www.gharchive.org/ , the overwhelming feeling I have is "creeped out". Taking periodic snapshots of repositories and their issues and wikis sounds good, but do we really need a log of every time someone watches an issue, and every commit message being irrevocably set on public record? That level of details on individual activity s…

If you put information online publicly, you should be always working under the assumption that it will be immediately archived by someone. Whether that is the Internet Archive for websites, a Discord bot archiving edits and deletions, Pushshift (formerly) for Reddit, or just some private group operating a web scraper.

At least in this situation the archived data is public.

Re: Getting 10TB of GitHub logs and extracting details of all users and repositories

#39

Earlier quoted context omitted.

Even https://archive.ph/atw1q didn’t work correctly on the page, it just ceases to scroll after a point.

The website works fine unless you enable javascript. That's usually the way it is with these sort of things. The webdev or CMS creates a perfectly functional website using HTML and CSS, then some javascript is added to shit the whole thing up. Disable javascript by default for a better web experience.

The noise is gone, enjoy your read! :)
Post reply on HN