Live data from Hacker News

Show HN: I scraped 3B Goodreads reviews to train a better recommendation model

book.sv

131–140 of 275 posts

Re: Show HN: I scraped 3B Goodreads reviews to train a better recommendation model

#132
post #130

I typed "Introduction to Real Analysis by Bartle" and I got: Steve Jobs by Walter Isaacson Harry Potter and the Sorcerer's Stone (Harry Potter, #1) by JK Rowling Topology by James R Munkres and so on.. Munkres' book is relevant and I want to read it, but what have Steve Jobs and Harry Potter got to do with with mathematics?

They have nothing to do with mathematics but everything to do with being extremely popular books.

Most people that have read a mathematics textbook have also read and enjoyed Harry Potter.

Given you have enjoyed drinking water and breathing in the past, there is a high likelihood that you will enjoy watching the Star Wars films.

Re: Show HN: I scraped 3B Goodreads reviews to train a better recommendation model

#135

Does this break part 4 of the Goodreads TOS? "[...] you agree not to sell, license, rent, modify, distribute, copy, reproduce, transmit, publicly display, publicly perform, publish, adapt, edit or create derivative works from any materials or content accessible on the Service. Use of the Goodreads Content or materials on the Service for any purpose not expressly permitted by this Agreement is strictly prohibited." Al…

Fairly meaningless in this day and age. Also IIRC scraping legality depends heavily on jurisdiction. Some places take a more permissive view of accessing publicly available information, even if a site's TOS forbids bots.

In the US there’s a major precedent [0] which held that scraping public-facing pages isn’t a CFAA "unauthorized access" issue. That’s a big part of why we’ve seen entire venture-backed scraping companies pop up - it’s not considered hacking if the data is already public.

[0] https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn

Re: Show HN: I scraped 3B Goodreads reviews to train a better recommendation model

#139
post #135

Does this break part 4 of the Goodreads TOS? "[...] you agree not to sell, license, rent, modify, distribute, copy, reproduce, transmit, publicly display, publicly perform, publish, adapt, edit or create derivative works from any materials or content accessible on the Service. Use of the Goodreads Content or materials on the Service for any purpose not expressly permitted by this Agreement is strictly prohibited." Al…

Fairly meaningless in this day and age. Also IIRC scraping legality depends heavily on jurisdiction. Some places take a more permissive view of accessing publicly available information, even if a site's TOS forbids bots. In the US there’s a major precedent [0] which held that scraping public-facing pages isn’t a CFAA "unauthorized access" issue. That’s a big part of why we’ve seen entire venture-backed scraping compa…

[deleted]
Post reply on HN