Bluesky Social Dataset (235M posts from 4M users)
1–10 of 46 posts
Re: Bluesky Social Dataset (235M posts from 4M users)
#2The dataset contains the complete post history of over 4M users (81% of all registered accounts), totaling 235M posts. We also make available social data covering follow, comment, repost, and quote interactions.
Since Bluesky allows users to create and bookmark feed generators (i.e., content recommendation algorithms), we also release the full output of several popular algorithms available on the platform, along with their timestamped “like” interactions and time of bookmarking.
This dataset allows unprecedented analysis of online behavior and human-machine engagement patterns. Notably, it provides ground-truth data for studying the effects of content exposure and self-selection, and performing content virality and diffusion analysis.
Re: Bluesky Social Dataset (235M posts from 4M users)
#3Re: Bluesky Social Dataset (235M posts from 4M users)
#4Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
Also, I feel like only recently there's been an influx of people who have actually interesting things to say so I'd love to see nextyear's dataset
Re: Bluesky Social Dataset (235M posts from 4M users)
#5associated paper: https://arxiv.org/abs/2404.18984
Re: Bluesky Social Dataset (235M posts from 4M users)
#6Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
I wonder how much time it takes to run this / what the script is / how resource intensive it is? Bsky is public right, so do you get rate limited? Do you scrape or use an official API? So many questions Also, I feel like only recently there's been an influx of people who have actually interesting things to say so I'd love to see nextyear's dataset
Re: Bluesky Social Dataset (235M posts from 4M users)
#7associated paper: https://arxiv.org/abs/2404.18984
hmm could you find the Github? I couldn't find it in the paper in the Code Availability section
The "code availability" says it's released "alongside [the dataset]", which appears to be the OP.
Re: Bluesky Social Dataset (235M posts from 4M users)
#8Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
Re: Bluesky Social Dataset (235M posts from 4M users)
#9Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
[deleted]
Re: Bluesky Social Dataset (235M posts from 4M users)
#10The trend toward everything being a walled garden is unfortunate.