Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
1–10 of 12 posts
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#2What are your impressions of scraper and html5ever? When I initially looked at HTML/XML parsing libraries for Rust, there didn't seem to be a standout library such as serde_json for JSON data. I was also considering using scraper + html5ever. However, I'm curious if scraper adds enough to warrant the additional dependency as opposed to directly using html5ever.
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#3Cool library, I might use this for a side project I have also parsing HTML data from HackerNews. What are your impressions of scraper and html5ever? When I initially looked at HTML/XML parsing libraries for Rust, there didn't seem to be a standout library such as serde_json for JSON data. I was also considering using scraper + html5ever. However, I'm curious if scraper adds enough to warrant the additional dependency…
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#4Cool library, I might use this for a side project I have also parsing HTML data from HackerNews. What are your impressions of scraper and html5ever? When I initially looked at HTML/XML parsing libraries for Rust, there didn't seem to be a standout library such as serde_json for JSON data. I was also considering using scraper + html5ever. However, I'm curious if scraper adds enough to warrant the additional dependency…
I haven't used scraper too much. I personally find the predicate approach of select.rs [0] easier to use. However in this case the selector approach just made more sense. Standalone html5ever can be a bit cumbersome to work with directly, scraper is basically an implementation of the html5ever's `TreeSink` trait, where as `select.rs` uses the hmtl5ever `RcDom` to parse the document but stores it in a more convenient…
Once I add the tokio feature, they all run as expected.
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#5Earlier quoted context omitted.
I haven't used scraper too much. I personally find the predicate approach of select.rs [0] easier to use. However in this case the selector approach just made more sense. Standalone html5ever can be a bit cumbersome to work with directly, scraper is basically an implementation of the html5ever's `TreeSink` trait, where as `select.rs` uses the hmtl5ever `RcDom` to parse the document but stores it in a more convenient…
Hey just FYI, when you run the hackernews and explore examples without enabling the tokio feature flag, you get compilation errors about undeclared types for all the tokio stuff. I think these examples just need entires in the Cargo.toml requiring the tokio feature like the reddit example. Once I add the tokio feature, they all run as expected.
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#6Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#7I tried to build a scraper in Rust just a few days ago and got stuck trying to concurrency limit my calls (the website I was scraping appropriately had a rate limit). I couldn’t figure out how to get tokio::stream/tokio_stream to work. Does this fix that problem?
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#8I tried to build a scraper in Rust just a few days ago and got stuck trying to concurrency limit my calls (the website I was scraping appropriately had a rate limit). I couldn’t figure out how to get tokio::stream/tokio_stream to work. Does this fix that problem?
Basically this a just a futures_timer::Delay [0] that is reset after each request, which is non blocking.
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#9I tried to build a scraper in Rust just a few days ago and got stuck trying to concurrency limit my calls (the website I was scraping appropriately had a rate limit). I couldn’t figure out how to get tokio::stream/tokio_stream to work. Does this fix that problem?
Yes, you can either limit how many request can be send concurrently or enforce a fixed or random delay in between consecutive requests. Basically this a just a futures_timer::Delay [0] that is reset after each request, which is non blocking. [0] https://docs.rs/futures-timer/3.0.2/futures_timer/
Re: Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
#10I tried to build a scraper in Rust just a few days ago and got stuck trying to concurrency limit my calls (the website I was scraping appropriately had a rate limit). I couldn’t figure out how to get tokio::stream/tokio_stream to work. Does this fix that problem?
It was recently used to implement rate limiting middleware for both tide[1] and actix[2]
[0] https://docs.rs/governor/0.3.1/governor/_guide/index.html