Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
[deleted]
Bluesky Social Dataset (235M posts from 4M users)
11–20 of 46 posts
Re: Bluesky Social Dataset (235M posts from 4M users)
#12Re: Bluesky Social Dataset (235M posts from 4M users)
#13Re: Bluesky Social Dataset (235M posts from 4M users)
#14Earlier quoted context omitted.
"Personal data" that was voluntarily published on a public microblogging platform with the explicit intention to share it with the world?
[deleted]
Re: Bluesky Social Dataset (235M posts from 4M users)
#15Re: Bluesky Social Dataset (235M posts from 4M users)
#16Re: Bluesky Social Dataset (235M posts from 4M users)
#17Re: Bluesky Social Dataset (235M posts from 4M users)
#18Re: Bluesky Social Dataset (235M posts from 4M users)
#19Earlier quoted context omitted.
[deleted]
"Personal data" that was voluntarily published on a public microblogging platform with the explicit intention to share it with the world?
Does Bluesky explicitly state the license the user will be publishing under (Creative Commons or whatever), or allow them to choose one?
Re: Bluesky Social Dataset (235M posts from 4M users)
#20Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-…
I wonder how much time it takes to run this / what the script is / how resource intensive it is? Bsky is public right, so do you get rate limited? Do you scrape or use an official API? So many questions Also, I feel like only recently there's been an influx of people who have actually interesting things to say so I'd love to see nextyear's dataset