Viewing profile — LisaG
LisaG
HN member- Joined
- Wed, Oct 20, 2010, 12:24 AM UTC
- HN karma
- 641
- Public activity
- 64 items
- HN profile
- View on Hacker News ↗
About LisaG
Recent public activity
-
comment
Comment #16845845
Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, h…
-
comment
Comment #16845815
Did you watch all of Hadley's video? You might get the title more if you saw/see the whole talk :)
-
comment
Comment #15195211
As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication t…
- story
-
comment
Comment #9814930
San Francisco CA Full-time / Onsite New, somewhat stealth startup, for profit company focused on social good. We have a very talented team so far comprised of : full stack web dev,…
-
comment
Comment #9814929
San Francisco CA Full-time / Onsite New, somewhat stealth startup, for profit company focused on social good. We have a very talented team so far comprised of : full stack web dev,…
-
comment
Comment #8150400
So excited so see Common Crawl data be useful for such fascinating work! I work at Common Crawl :)
-
comment
Comment #8076014
Love this idea!!
-
comment
Comment #7518222
Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more t…
- story
- story
- story
-
comment
Comment #6936720
I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much. This new version is a whole different animal. Not only is it much prettier (…
-
comment
Comment #6811843
There will be news about a subset sometime next month!
- story
-
comment
Comment #6214957
If you don't feel like reading the paper Sebastian wrote on the Common Crawl data, he gives a summary of his findings in this video. Link to full paper: http://bit.ly/14dxSJq
- story
-
comment
Comment #6209678
We do think it is worth it to avoid duplicative efforts. Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data …
-
comment
Comment #6209647
Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on variou…
-
comment
Comment #6208864
Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl take…
- story
- comment
-
comment
Comment #6114351
That's interesting. We also could revive phage therapy which uses bacteriophages (viruses that replicate in bacteria). http://en.wikipedia.org/wiki/Phage_therapy
-
comment
Comment #5505259
If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.
-
comment
Comment #5327850
I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code. If you didn't see the details…