Live data from Hacker News

Viewing profile — LisaG

LisaG

HN member
Joined
Wed, Oct 20, 2010, 12:24 AM UTC
HN karma
641
Public activity
64 items

About LisaG

Science geek with a strong affinity for computer nerds. Director of Common Crawl www.commoncrawl.org Former Chief of Staff at Creative Commons www.creativecommons.org

Recent public activity

  1. comment
    Comment #16845845

    Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, h…

  2. comment
    Comment #16845815

    Did you watch all of Hadley's video? You might get the title more if you saw/see the whole talk :)

  3. comment
    Comment #15195211

    As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication t…

  4. story
  5. comment
    Comment #9814930

    San Francisco CA Full-time / Onsite New, somewhat stealth startup, for profit company focused on social good. We have a very talented team so far comprised of : full stack web dev,…

  6. comment
    Comment #9814929

    San Francisco CA Full-time / Onsite New, somewhat stealth startup, for profit company focused on social good. We have a very talented team so far comprised of : full stack web dev,…

  7. comment
    Comment #8150400

    So excited so see Common Crawl data be useful for such fascinating work! I work at Common Crawl :)

  8. comment
    Comment #8076014

    Love this idea!!

  9. comment
    Comment #7518222

    Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more t…

  10. story
  11. story
  12. story
  13. comment
    Comment #6936720

    I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much. This new version is a whole different animal. Not only is it much prettier (…

  14. comment
    Comment #6811843

    There will be news about a subset sometime next month!

  15. story
  16. comment
    Comment #6214957

    If you don't feel like reading the paper Sebastian wrote on the Common Crawl data, he gives a summary of his findings in this video. Link to full paper: http://bit.ly/14dxSJq

  17. story
  18. comment
    Comment #6209678

    We do think it is worth it to avoid duplicative efforts. Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data …

  19. comment
    Comment #6209647

    Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on variou…

  20. comment
    Comment #6208864

    Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl take…

  21. story
  22. comment
  23. comment
    Comment #6114351

    That's interesting. We also could revive phage therapy which uses bacteriophages (viruses that replicate in bacteria). http://en.wikipedia.org/wiki/Phage_therapy

  24. comment
    Comment #5505259

    If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.

  25. comment
    Comment #5327850

    I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code. If you didn't see the details…