Live data from Hacker News

Show HN: Crawlee – Web scraping and browser automation library for Node.js

crawlee.dev

81–86 of 86 posts

Re: Show HN: Crawlee – Web scraping and browser automation library for Node.js

#81
post #58

Earlier quoted context omitted.

people trash JavaScript while using python? They should look at the mirror! Anyone can trash js except for all these C like languages with boringly similar designs, doubly so for python. Itself took over the world due to the fact that amateur (scientists and other professions as opposed to programmers) can easily play with it.

Python is anything but a C-like language, I really don't know what you're talking about.

Because of indentation? Gimme a break, most of the semantics are the same. Different would be like lisp or SML

Re: Show HN: Crawlee – Web scraping and browser automation library for Node.js

#82
post #74
post #72

Oh this looks lovely, congratulations! I would really like this but running in Python.

you want something better than Scapy? maybe I have Stockholm syndrome but I find it to be very well structured, testable, and has solved every problem I've had with running scrapers

Do you feel this and Scrapy are similar? In my reading this has a different feature set.

It allows headed crawling + avoiding blockers etc.

Re: Show HN: Crawlee – Web scraping and browser automation library for Node.js

#84
post #82
post #74

Earlier quoted context omitted.

you want something better than Scapy? maybe I have Stockholm syndrome but I find it to be very well structured, testable, and has solved every problem I've had with running scrapers

Do you feel this and Scrapy are similar? In my reading this has a different feature set. It allows headed crawling + avoiding blockers etc.

In that they're trying to be crawling frameworks, and for sure Scrapy allows headed crawling via Splash, it's just not something I've needed or advocate for

Scrapy also has a long lineage of extensions, which maybe Crawlee will gain as it increases in popularity but I didn't see any obvious way of decoupling (for example) if one wanted to plug a new storage engine into Crawlee: https://crawlee.dev/docs/guides/result-storage versus that delineation is very strong in Scrapy for all its moving parts

Also, Parsel (the selector library powering Scrapy) is A++, in that it allows expressing one's intent via xpath, css selector, and regex matches in a fluent API; I'm sure this nodejs framework allows doing something similar because it seems to be all-in on the DOM, but it for sure will not be `response.xpath("//whatever").css("#some-id").re("firstName: (.+)").extract()`

Further, as I mentioned -- and as someone pointed out elsewhere in this submission -- Scrapy is prepared to store requests to disk and makes testing spider methods super easily since they're very well defined callback methods. If you have the HTML from a prior run and need to reproduce a bad outcome, testing just the "def parse_details_page" is painless. It certainly may be possible to test Crawlee code, too, but I didn't see anything mentioned about it

Re: Show HN: Crawlee – Web scraping and browser automation library for Node.js

#85
post #84
post #82

Earlier quoted context omitted.

Do you feel this and Scrapy are similar? In my reading this has a different feature set. It allows headed crawling + avoiding blockers etc.

In that they're trying to be crawling frameworks, and for sure Scrapy allows headed crawling via Splash, it's just not something I've needed or advocate for Scrapy also has a long lineage of extensions, which maybe Crawlee will gain as it increases in popularity but I didn't see any obvious way of decoupling (for example) if one wanted to plug a new storage engine into Crawlee: https://crawlee.dev/docs/guides/result-…

so what would be the the approach with this library? used scrapy and like it but more in the JS ecosystem now so would like this to be similair.

Re: Show HN: Crawlee – Web scraping and browser automation library for Node.js

#86
post #85
post #84

Earlier quoted context omitted.

In that they're trying to be crawling frameworks, and for sure Scrapy allows headed crawling via Splash, it's just not something I've needed or advocate for Scrapy also has a long lineage of extensions, which maybe Crawlee will gain as it increases in popularity but I didn't see any obvious way of decoupling (for example) if one wanted to plug a new storage engine into Crawlee: https://crawlee.dev/docs/guides/result-…

so what would be the the approach with this library? used scrapy and like it but more in the JS ecosystem now so would like this to be similair.

I don't know anything about this other than the announcement here and on reddit, so you'll likely want to post your question as a top comment so Jan can see it, or open a GH issue so they can help you evaluate
Post reply on HN