Hey there, first off cool stuff I can see this being plenty useful for various projects. We're exploring GDELT data for our own needs and I was wondering if you wouldn't mind sharing what were some rough spots or gotchas using the project?
Thanks! There are a lot of issues with GDELT data, but the things that come to mind recently are: - Missing keywords. Huge topics like "brexit" won't be found, so you need to extract those yourself. - Dirty titles. You need to extract the article titles from the website metadata, and more often than not they'll pre/append their site name like "[title] - DailyStormer" and links like "[title]: Business, Stocks, News" e…
https://www.splcenter.org/fighting-hate/extremist-files/indi...