Live data from Hacker News

I scraped 1.94M Airbnb photos for opium dens, pet cameos, and messy kitchens

burla-cloud.github.io

1–10 of 44 posts

Re: I scraped 1.94M Airbnb photos for opium dens, pet cameos, and messy kitchens

#6
"Looking at every public Airbnb listing in Inside Airbnb's open data dump, all at once, on Burla"

This Inside Airbnb?

Community Guidelines

Please:

Only take the data you need. Do not scrape data from the site, if you would like to subscribe to the data directly, please email data@insideairbnb.com

Re: I scraped 1.94M Airbnb photos for opium dens, pet cameos, and messy kitchens

#7
post #6

"Looking at every public Airbnb listing in Inside Airbnb's open data dump, all at once, on Burla" This Inside Airbnb? Community Guidelines Please: Only take the data you need. Do not scrape data from the site, if you would like to subscribe to the data directly, please email data@insideairbnb.com

>Everything was parallelized on Burla, on a single dynamic cluster that scaled to ~1.7K CPU workers for photo download and CLIP, with 20 A100 GPUs running embedding clusters in parallel on the same cluster.

That's a lot of budget - would have been nice if they'd made an actual donation to the project, instead of pounding the project's servers and bandwidth when there are much better ways to interact with the data.

Re: I scraped 1.94M Airbnb photos for opium dens, pet cameos, and messy kitchens

#9
What a waste of energy (money/resources)... Scraping and AI-scanning 2 million photos to identify animals in the advertisement pictures? What's the point.

As an exercise a sample of 1000 photos would've been enough. As a database, knowing a listing has a cat in the picture or a funny review doesn't offer any real value.

I wonder what the footprint is of such an exercise.

Post reply on HN