Live data from Hacker News

Viewing profile — mlucy

mlucy

HN member
Joined
Sun, May 17, 2015, 5:06 AM UTC
HN karma
91
Public activity
34 items

About mlucy

I'm currently working on https://www.basilica.ai , an API that makes it easy for ML teams to work with high-dimensional data like images and text.

Recent public activity

  1. comment
    Comment #25652878

    This is really cool! I especially like that there's a premade colab notebook that lets you play with it: https://colab.research.google.com/github/openai/clip/blob/ma... . I'm a lit…

  2. comment
    Comment #19165746

    I don't think you could get access to the actual models that are being used to run e.g. Google Translate, but if you just want a big pretrained model as a starting point, their res…

  3. comment
    Comment #19165303

    > Is it always good enough to take the outputs of the next-to-last layer as features? It usually doesn't matter all that much whether you take the next-to-last or the third from la…

  4. comment
    Comment #19165225

    There's actually been a lot of really good work recently around textual transfer learning. Google's BERT paper does sentence-level pretraining and transfer to get state of the art …

  5. comment
    Comment #19165200

    Yeah, I think this pattern is pretty common. (Basilica's main business is an API that does deep feature extraction as a service, so we end up talking to a lot of people with tasks …

  6. comment
    Comment #19164985

    I think if you have a small to medium sized dataset of images or text, deep feature extraction would be the first thing I'd try. I'm not sure what the most interesting problems wit…

  7. comment
    Comment #19164916

    I hadn't read it before! That's a fascinating result, actually. They emphasize interpretability in the paper, but I find it more interesting that you can do so well with only local…

  8. comment
    Comment #19164826

    I don't work with time series data much myself. I would imagine you can get at least some transfer learning, since there are patterns that show up across different domains. It look…

  9. comment
    Comment #19164697

    Definitely. There's been a lot of exciting work recently for text in particular, like https://arxiv.org/pdf/1810.04805.pdf .

  10. comment
    Comment #19164674

    Linear Algebra Done Right would be my recommendation.

  11. comment
    Comment #19164484

    Hi everyone! Author here. Let me know if you have any questions, this is one of my favorite subjects in the world to talk about.

  12. comment
    Comment #19164481

    > He's almost describing a future where we might buy/license pre-trained models from Google/Facebook/etc that are trained on huge datasets, and then extend that with more specific …

  13. comment
    Comment #18982255

    Really cool idea. I hope you manage to get into a sustainable cycle of people you've helped with bankruptcy getting back on their feet and donating to help others in the same posit…

  14. comment
    Comment #18914427

    I would second this; sentence embeddings outperform word embeddings on basically all tasks where you actually have sentences to work with. The only downside is that they're signifi…

  15. comment
    Comment #18914381

    Interesting. It took me a while to figure out what the main contribution here is, since doing dimensionality reduction on embeddings is fairly common. I think the main contribution…

  16. comment
    Comment #18869827

    A word embedding transforms a word into a series of numbers, with the property that similar words (e.g. "dog" and "canine") produce similar numbers. You can have embeddings for oth…

  17. comment
    Comment #18869184

    It's really difficult to overstate how important embeddings are going to be for ML. Word embeddings have already transformed NLP. Most people I know, when they sit down to work on …

  18. comment
    Comment #18351455

    Hi there :) Apologies for the super long response, but you had a lot of points. > Am I really missing something here or this thing is a complete nonsense with no actual use cases w…

  19. comment
    Comment #18350021

    You can definitely improve performance by choosing an embedding closely related to your task. In the future we're hoping to have more embeddings for specialized tasks. Kind of surp…

  20. comment
    Comment #18349948

    Thanks! No production use cases yet. This is the first usable release, and it's the bare minimum we felt we could build before showing it to people. > Are you a YC company? We have…

  21. comment
    Comment #18349912

    We aren't currently doing this. In the future I think we'll try to embed into a single space on a best-effort basis, assuming we can find the engineering resources. It will be real…

  22. comment
    Comment #18349304

    That's a really interesting idea. I can't really think of a barrier to this. Detecting the file format is straightforward, and generic image/text/etc. embeddings work surprisingly …

  23. comment
    Comment #18349193

    Hey! We're embedding images by feeding them through a deep neural net and using the activations of an intermediate layer as an embedding. You can read https://arxiv.org/abs/1403.63…

  24. comment
    Comment #18348849

    "Word2vec for anything" is where we want to get to. Right now we only support images and text, but you can see the other data types on our roadmap at https://www.basilica.ai/availa…

  25. comment
    Comment #18348305

    Yeah, I agree. The number of new nouns per year is kind of ridiculous.