Live data from Hacker News

Bluesky Social Dataset (235M posts from 4M users)

zenodo.org

21–30 of 46 posts

Re: Bluesky Social Dataset (235M posts from 4M users)

#21
post #5

Earlier quoted context omitted.

hmm could you find the Github? I couldn't find it in the paper in the Code Availability section

Is it "scripts.tar.gz. A collection of Python scripts, including the ones originally used to crawl the data, and to perform experiments. These scripts are detailed in a document released within the folder" in the OP? The "code availability" says it's released "alongside [the dataset]", which appears to be the OP.

oh good eye, I didn't catch that

Re: Bluesky Social Dataset (235M posts from 4M users)

#23
post #9

Earlier quoted context omitted.

"Personal data" that was voluntarily published on a public microblogging platform with the explicit intention to share it with the world?

Would it make a difference if we were talking about articles on a news website? I'm kind of on the fence on this one but I can see the point of view that just posting something online doesn't necessarily grant the end user an unlimited license to use the data. Source code is another example; open-sourcing a project doesn't automatically give someone else the right to use that code in their own projects. Does Bluesky…

> Would it make a difference if we were talking about articles on a news website.

News articles are pretty explicitly copyrighted and published for a commercial purpose. The websites make their terms clear when you visit. I don't think anyone can argue that it is legal to copy and distribute these articles, same as a book or movie or song.

Data posted on Bluesky on the other hand is meant to be broadly shared using the AT protocol. It is quite literally a feature. If you create your own Bluesky client, for example, you aren't committing copyright violation by downloading someone else's posts on there. Similarly, you aren't going against any terms of service by consuming a firehose of data from an AT relay.

Re: Bluesky Social Dataset (235M posts from 4M users)

#24
post #23

Earlier quoted context omitted.

Would it make a difference if we were talking about articles on a news website? I'm kind of on the fence on this one but I can see the point of view that just posting something online doesn't necessarily grant the end user an unlimited license to use the data. Source code is another example; open-sourcing a project doesn't automatically give someone else the right to use that code in their own projects. Does Bluesky…

> Would it make a difference if we were talking about articles on a news website. News articles are pretty explicitly copyrighted and published for a commercial purpose. The websites make their terms clear when you visit. I don't think anyone can argue that it is legal to copy and distribute these articles, same as a book or movie or song. Data posted on Bluesky on the other hand is meant to be broadly shared using t…

Right, that's why I asked about Bluesky's content license; just because it's not in your face when you visit, doesn't mean you don't have to abide by it.

You understand that categories of usage are important, right? No-one is breaking the GPL by reading source code, but incorporating into your own codebase can be problematic if not done correctly. Similarly, human beings reading the data posted by a Bluesky user is not the same as aggregating and analysing the data of thousands of users. As I said I'm on the fence with this, but I do understand why someone might have a problem with it.

Re: Bluesky Social Dataset (235M posts from 4M users)

#26

I’m glad to see a new platform that isn’t completely locked down, allowing analysis like this. The trend toward everything being a walled garden is unfortunate.

[flagged]

Actually, since this isn’t locked up by the big copyright holders, we can all use it and profit.

Re: Bluesky Social Dataset (235M posts from 4M users)

#27

I’m glad to see a new platform that isn’t completely locked down, allowing analysis like this. The trend toward everything being a walled garden is unfortunate.

[flagged]

Not all the profit, really. All would imply there was no value to begin with. I get the dislike, but i still comment on the open web because it has value to me. I'm still willing to answer questions on SO/reddit/etc because it has value to at least one (and hopefully more) people. That hasn't changed.

Not sure what to say about companies making money off of my data.. but the posting itself doesn't seem to be that much of a negative.

Thoughts? I see this sentiment a lot and it almost feels like "open" is bad these days. If anything i feel it almost is more important than ever.. as we're on the cusp of no need to ever go to forums/interact/etc.

Re: Bluesky Social Dataset (235M posts from 4M users)

#28
post #6
post #4

Earlier quoted context omitted.

I wonder how much time it takes to run this / what the script is / how resource intensive it is? Bsky is public right, so do you get rate limited? Do you scrape or use an official API? So many questions Also, I feel like only recently there's been an influx of people who have actually interesting things to say so I'd love to see nextyear's dataset

Not sure about bulk export but you can set up a full stream of all activity without even registering an account.

Blows my mind that they can send that much for free.

Re: Bluesky Social Dataset (235M posts from 4M users)

#29
post #26

Earlier quoted context omitted.

[flagged]

Actually, since this isn’t locked up by the big copyright holders, we can all use it and profit.

How will you use it to profit? You don’t have sweetheart cloud deals on ML training clusters. This benefits big players, not us.

Re: Bluesky Social Dataset (235M posts from 4M users)

#30
post #29
post #26

Earlier quoted context omitted.

Actually, since this isn’t locked up by the big copyright holders, we can all use it and profit.

How will you use it to profit? You don’t have sweetheart cloud deals on ML training clusters. This benefits big players, not us.

Most impactful ML can be created on colab. Not Chatgpt, but most of the stuff not on the long tail.
Post reply on HN