Live data from Hacker News

Launch HN: Superb AI (YC W19) – AI-Powered Training Data

news.ycombinator.com

11–20 of 25 posts

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#11

One of the more interesting startups to come out of W19 in my opinion. How big is your human-in-the-loop piece? Staff size?

Thank you so much! The human-in-the-loop piece really starts to kick in for larger volumes of data where we can iterate the fine-tuning process several times. As an example, we were able to achieve +30% speed boost for a client after a few cycles of the loop. We are a team of 13, including 9 engineers.

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#12

If your AI can accurately label my training data, why wouldn’t I just use your AI for my application?

Hi, I'm Jonghyuk, one of the co-founders. One point I would like to add is that a 90% accurate AI model may not be very useful for an application, but with the right data pipeline and well-designed system, we can extract quite a bit of boost out of it for data annotation.

Numerically, how much is "quite a bit of a boost"?

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#13

Earlier quoted context omitted.

Hi, I'm Jonghyuk, one of the co-founders. One point I would like to add is that a 90% accurate AI model may not be very useful for an application, but with the right data pipeline and well-designed system, we can extract quite a bit of boost out of it for data annotation.

Numerically, how much is "quite a bit of a boost"?

It depends on how accurate our AI performs on a particular task, but as a back-of-the-envelope calculation, if we had a 90% accurate AI that means human annotators only have to work on the remaining 10%, giving us 10x boost. Obviously, there is some overhead not accounted for in this calculation, but with our current technology we can boost up to 10x depending on the type of data.

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#14

Speaking from my previous experience of running a software testing company, I am skeptical about scaling In-house professional team. Is In-house workforce cost effective ?

Thanks for pointing this out. Although using crowdsourced labor could be the most cost-effective, we don't think it can guarantee the level of quality we get from in-house team. We believe our AI will be able to automate a lot of the pieces and make our in-house team cost-effective. I'm sure you have a lot of experience in terms of scaling workforce, please let me know if you have any advice for us!

Bespoke, artisanal data.

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#15

Earlier quoted context omitted.

Numerically, how much is "quite a bit of a boost"?

It depends on how accurate our AI performs on a particular task, but as a back-of-the-envelope calculation, if we had a 90% accurate AI that means human annotators only have to work on the remaining 10%, giving us 10x boost. Obviously, there is some overhead not accounted for in this calculation, but with our current technology we can boost up to 10x depending on the type of data.

How do you know which are the 90% it got wrong and which is the 10% it got right?

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#16
post #2

Congratulations on your launch! I think this space is going to be super interesting and allowing companies to build products like this by simply adding an API to their pipeline is awesome. I had a question, how would you point out the differences between you and Scale API[0]? [0]: https://scale.ai

Interesting. I didn't know about scale. How would someone figure out that these tools are needed by so and so companies => let's provide APIs. I was of the idea that all the self driving startups have their own people solving these challenges. Why would anyone use scale?

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#18
post #15

Earlier quoted context omitted.

It depends on how accurate our AI performs on a particular task, but as a back-of-the-envelope calculation, if we had a 90% accurate AI that means human annotators only have to work on the remaining 10%, giving us 10x boost. Obviously, there is some overhead not accounted for in this calculation, but with our current technology we can boost up to 10x depending on the type of data.

How do you know which are the 90% it got wrong and which is the 10% it got right?

We have both AI-assisted and manual inspections in the pipeline. A good analogy would be an assembly line where humans and machines collaborate not only for building things but also for the quality control (ie. vision inspection system + manual inspection)

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#19
post #2

Congratulations on your launch! I think this space is going to be super interesting and allowing companies to build products like this by simply adding an API to their pipeline is awesome. I had a question, how would you point out the differences between you and Scale API[0]? [0]: https://scale.ai

Interesting. I didn't know about scale. How would someone figure out that these tools are needed by so and so companies => let's provide APIs. I was of the idea that all the self driving startups have their own people solving these challenges. Why would anyone use scale?

It's really a matter of productivity. Creating training sets of data is a labor intensive process. It may not be worth the cost to do so, if they can outsource that work instead.

Re: Launch HN: Superb AI (YC W19) – AI-Powered Training Data

#20

This is great. Thank you for posting. What humans do you have tagging the images, after the AI portion? Are they the 13 employees you have now? Do you think that you will need to focus on a couple of verticals, so that your AI will have more of an impact?

Thank you for the interest! We have a team of ~50 in-house annotators that use our AI tools and we are looking to expand out to Southeast Asia, namely Vietnam or the Philippines, to set up our annotator workforce there in the near future.

As you mentioned, we see that some data labeling companies focus on a few verticals like autonomous vehicles. These verticals are very data hungry and we actually do have clients in these sectors, but I also see a huge opportunity in the less "AI-savvy" verticals such as consumer electronics, physical security, factory automation and so on. They not only have a huge need for AI, but they also lack the AI talent to build AI themselves. So, one natural possibility is that we extend our service and actually deliver the AI built on top of the training data we make. To do that, we will need to automate that piece as well using AutoML or Meta-Learning which the co-founders already have experience with (AI building AI!). It's also possible that we stay focused on just the training data piece for a few of these verticals.

Post reply on HN