Live data from Hacker News

Viewing profile — pgao

pgao

HN member
Joined
Sun, Mar 11, 2012, 11:14 PM UTC
HN karma
60
Public activity
28 items

About pgao

No profile information was provided.

Recent public activity

  1. story
    Show HN: Tidepool – analytics for large text datasets

    Hello HN! I'm Peter, one of the folks who helped create Tidepool. We last shared Tidepool with HN about 7 months ago https://news.ycombinator.com/item?id=36957762 Since then, the A…

  2. comment
    Comment #36957763

    With the last few months, there's been a Cambrian explosion of products integrating AI. A big part of this is because LLMs enable users to get good performance on a lot of NLP task…

  3. story
  4. story
  5. comment
    Comment #34912646

    From previous experience doing applied ML, it's pretty hard to do pretty basic operations on data - splitting broad classes into fine classes, collecting data that the model strugg…

  6. story
  7. story
  8. job
  9. comment
    Comment #29083545

    Aquarium ( https://aquariumlearning.com/ ) | Remote Only (North American Timezones) | Full Time Aquarium is an ML data management system that helps ML teams improve their models by…

  10. story
  11. story
  12. comment
    Comment #26031196

    I think active learning has a time and a place. If you're getting started with a project from scratch, you probably don't need active learning for the exact reasons you describe - …

  13. comment
    Comment #26031148

    Former self-driving engineer here. I'm also pretty skeptical about synthetic data. For the scenario you described, it turns out that if you drive enough, you'll eventually see some…

  14. comment
    Comment #26031123

    Here's a great paper to get started with! https://arxiv.org/abs/1703.04977

  15. story
  16. story
  17. story
  18. story
  19. comment
    Comment #23826571

    Yup, I saw your form submission through our site! I reached out to you over email, I'm confident we can help out :)

  20. comment
    Comment #23826534

    Thanks! We don't expect users to upload all data to our service - the type of data we're interested in is "metadata." URLs to the raw data, labels, inferences, embeddings, and any …

  21. comment
    Comment #23824612

    Thanks for the shoutout! We got connected to jononor through our previous r/machinelearning launch: https://www.reddit.com/r/MachineLearning/comments/hjbl4h/p_l...

  22. comment
    Comment #23824599

    Thanks a bunch! I think the biggest issues with this approach is the requirement for embeddings. It's hard sometimes for a customer to understand what layer to pull out of their ne…

  23. comment
    Comment #23824526

    I absolutely, 120% agree on the importance of adding the right data. Aquarium helps you with: "what data should I be collecting to improve my model" and "where do I find that data?…

  24. comment
    Comment #23824064

    Yes, your interpretation is correct. I don't think we're going to get into synthetic data generation in the near term, mainly due to the amount of effort required + questions about…

  25. comment
    Comment #23821880

    Hey there, it's a product right now! Our goal is to make it self serve, but we're currently onboarding people one-by-one manually until we can streamline the onboarding flow and bu…