Live data from Hacker News

Show HN: my weekend project – browse jobs posted to HN

hackernewsjob.com

11–20 of 31 posts

Re: Show HN: my weekend project – browse jobs posted to HN

#11
post #8

I wish there could be filters. I could sort by city, type of job, pay etc. But a good start.

Thanks. I thought about adding that. The posting text is pretty unstructured so I thought a search would get 80% of the way there. Tricky to extract city & type & pay from what is basically a wall of text. Before this I was doing a bunch of ctrl+f next next next searches on the original content. One thing you can do is click the little 'x' to right of a post to hide any post you're not interested in or already contacted. That will persist until you clear your cookies so you can kinda personalize it a bit.

Re: Show HN: my weekend project – browse jobs posted to HN

#14
post #8

I wish there could be filters. I could sort by city, type of job, pay etc. But a good start.

Thanks. I thought about adding that. The posting text is pretty unstructured so I thought a search would get 80% of the way there. Tricky to extract city & type & pay from what is basically a wall of text. Before this I was doing a bunch of ctrl+f next next next searches on the original content. One thing you can do is click the little 'x' to right of a post to hide any post you're not interested in or already contac…

There are a few standard things you could look for, like two letters for state or SF or NY{,C}.

Re: Show HN: my weekend project – browse jobs posted to HN

#15

Just curious what technologies you are using? Also how quickly will this update on the 1st when a new thread is posted...again - just curious. Great job! Much better than ctrl-f-ing my way through postings [just like you were doing].

Thanks.

To answer your first question:

I'm using a custom scraper to get the results. I request the job pages through a random http proxy every so often, parse them, and update the cached json of all the comments. Getting a reliable, free, machine-friendly list of http proxies is harder than I thought. It was automatic, but the proxy provider I was using went down (doh!) so now I manually pull one from a list and trigger the scrape.

On the server I'm using a tiny express app (node.js) + redis to store hidden posts per use based on a cookie.

On the browser just a little bit of twitter boostrap + jquery + some custom UI helpers to generate the page. The searching and hide/show is done within the browser itself. I try to involve the server as little as possible.

To answer your second question:

It will update as soon as I kick off another scrape. I hope to re-enable the automation on this soon so it will only lag 10-15 minutes behind the actual page. Right now it could be quite a while (up to 12 hours).

Yeah! ctrl-f was tough. Especially as the comments spanned multiple pages and weren't sortable by date.

Post reply on HN