Live data from Hacker News

Viewing profile — stephensonsco

stephensonsco

HN member
Joined
Sun, Sep 13, 2015, 3:00 AM UTC
HN karma
419
Public activity
112 items

About stephensonsco

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #18711726

    I didn't answer "Do you have a cheap way of generating high quality data?". We have good ways to do it. They're not that cheap though. It's expensive (organizationally and real $$$…

  3. comment
    Comment #18711689

    Yep, all good points. One thing to consider is that generalization is a big problem. It's easy to get good on a specific dataset nowadays (like 5-10% word error rate level on acade…

  4. comment
    Comment #18711650

    You can pay to get training at a lower usage amount too.

  5. comment
    Comment #18711444

    You find out very interesting things even randomly sampling your life in audio. We still come back to this for fun. The original device was an intel edison but recent variants have…

  6. comment
    Comment #18711428

    If you're training from scratch around 10k hours is needed to get a good model, but when you are transfer learning you don't need nearly that much (100 hours gets you a lot). We ex…

  7. comment
    Comment #18711382

    There's no additional charge for training a custom model when your usage is a minimum of 10k hrs/mo.

  8. comment
    Comment #18710517

    Best to say "yes! but only some of the time". It's something we're working on right now. You can be 80% accurate, by some metric, but it's still not good enough usually to pass a h…

  9. comment
    Comment #18710429

    We do custom models (train the full DNN, not just tack on a new text language model) using transfer learning and it works for small numbers of examples too. Glad to hear you asking…

  10. comment
    Comment #18710237

    It's a metric that's hard to nail down because there is so much parameter space that you are flattening into one number. Also it doesn't address the "I care about these five high v…

  11. comment
  12. comment
    Comment #18709882

    This is a seriously fertile area where you get to "define the new interface". It's a big problem though, since few buyers know they want those things. Around 95% of customers come …

  13. comment
    Comment #18709745

    Price starts at $1/hr billed in 1 second increments. Frequently we charge less than that, since the price is dropped with volume, and that's typically businesses have a steady amou…

  14. story
    Launch HN: Deepgram (YC W16) – Scalable Speech API for Businesses

    Hey HN, I’m Scott Stephenson, one of the cofounders of Deepgram ( https://www.deepgram.com/ ). Getting information from recorded phone calls and meetings is time-intensive, costly,…

  15. comment
    Comment #17989982

    Pretrained models are very nice to spin up a system and start using it. You need that because training the model is so hard. But, the pretrained models are by no means a good gener…

  16. comment
    Comment #17989918

    You're doing awesome (arduous) work. The text normalization is especially a total bear. I feel your pain. Limiting your text to one file is good in many ways because it allows you …

  17. story
  18. comment
    Comment #17218009

    Anyone able to find any data on time 'til accuracy? I don't see it (even in the video linked in another comment). Sure it's nice to "achieve 90% scaling efficiency" in images/sec, …

  19. comment
    Comment #16827450

    The dark matter is a gas not a solid. It's not orbiting the sun, it's orbiting the galaxy (technically, not even the galaxy, it's orbiting inside it's own halo mostly). The orbits …

  20. comment
    Comment #16824410

    Also very true. Unfortunately this data is dirty, context dependent, or plain missing. This is true of pretty much any "real" experiment.

  21. comment
    Comment #16824388

    This is awesome. But it is still impractical. The researchers knee deep in this already have to fight for the computing power and data access (and error checking of all of that) fo…

  22. comment
    Comment #16824320

    In the direct-search-for-dark-matter-experimental-community the DAMA results have long been excluded and therefore ~discredited (see lots of papers from LUX/XENON/PANDAX/CDMS/etc).…

  23. comment
    Comment #16824262

    So you think the LHC should "publish" 100 petabytes of data? What you are stating isn't practical because of cost. But I'll definitely go with the idea that "if you want to make a …

  24. comment
    Comment #16371636

    The problem is that a founder who hasn't already filled that role themselves doesn't know two things: 1) what the key parts of doing the job are and 2) which skills/personality the…

  25. comment
    Comment #16109635

    build it, you'll have a hit