Live data from Hacker News

Launch HN: Datasaur (YC W20) – data labeling interface for NLP

news.ycombinator.com

21–30 of 67 posts

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#21
post #15

Congratulation for the launch! To understand the scope of your work a little bit, if I have Prodigy with custom labeling needs set up for me, do I still benefit from switching to datasaur?

Apologies for the delay! There is some overlap with what Prodigy works on and I'm a big fan of what they're working on. We cover some additional use cases (like coreference parsing) and additionally help with managing teams of labelers. We're complementary in many regards. Happy to discuss further, based on your labeling needs.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#22

On the pricing page, the Growth box shows a checkmark for "Unlimited labels" but right below in the "Choose the right plan for you", the Growth plan says the number of labels is 10,000,000.

Great catch! We'll correct it asap. Since you caught it, we'll give you unlimited labels :)

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#23

Interesting product. Could have used this at previous companies. How is this different from FigureEight or Scale?

Scale offers labeling as a service. Datasaur is an interface that companies can buy for their own labeling personnel, if I understand correctly.

That's right! Scale probably has some awesome internal tools that help them label faster. Datasaur wants to make those same optimizations available to anyone with their own labelers.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#24
Could you please elaborate on what you mean by "intelligently validate the quality of labels in a document and complement human judgment", and discuss your methodology?

This seems to operate under the assumption that human labels are not actually the ground truth. I understand that they can be dirty, but most unsupervised approaches aren't producing a ground truth, either. So, are you saying it's better to have multiple pretty good sources of truth instead? Because depending on the application, that might make sense or it might be like trying to start a farm with a dead horse and a dead cow.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#25
Wish you good luck, the website looks clean, the product idea is good:) You request an image however which width is 3000+ pixels: https://s.datasaur.ai/static/media/homepage-hero.4917b8af.pn... . 1200px in width should be enough, I would resize the image, it slows down the page.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#26
post #25

Wish you good luck, the website looks clean, the product idea is good:) You request an image however which width is 3000+ pixels: https://s.datasaur.ai/static/media/homepage-hero.4917b8af.pn... . 1200px in width should be enough, I would resize the image, it slows down the page.

Yikes - good point. We'll optimize.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#28

Could you please elaborate on what you mean by "intelligently validate the quality of labels in a document and complement human judgment", and discuss your methodology? This seems to operate under the assumption that human labels are not actually the ground truth. I understand that they can be dirty, but most unsupervised approaches aren't producing a ground truth, either. So, are you saying it's better to have multi…

Certainly. Our philosophy is to complement human wisdom with computer precision. Humans may often be labeling for 8 hours a day and may get fatigued. So if Starbucks has been labeled as a cafe 35x in a document and as a person 2x, we can flag this and ask "hey, are you sure you wanted to label this as a person?". Or if we know for a fact Canada is a country, but it's labeled as an animal in a document, we can raise a flag as well. This won't work for everything, but we think it can help with quality assurance.

Re: Launch HN: Datasaur (YC W20) – data labeling interface for NLP

#30
LinkedIn suggested a post from you a couple weeks ago and I remember thinking “what’s Ivan up to?” and I saw Datasaur. Congrats on YC! I know that our time at Yahoo was a brief overlap but I remember the swirl of ML, Knowledge Graph and labelling our org was at 5 years ago.

Good luck with Datasaur!

- Julio Nobrega

Post reply on HN