Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
11–20 of 60 posts
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#12I saw your recent $24M series A and was kind of surprised to only see you launching now, congrats! YC seems to fund quite many document extraction companies, even within the same batch: - Pulse (YC W24): https://www.ycombinator.com/companies/pulse-3 - OmniAI (YC W24): https://www.ycombinator.com/companies/omniai - Extend (YC W23): https://www.ycombinator.com/companies/extend How do you differentiate from these? And h…
I assume y'all launched before this to select partners? Or perhaps this is a new product on top of the core product?
Congrats! Keep at it!
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#13I'm not a product fit, but I would like to take a moment to praise the detailed beauty of the design work on the site. From the typography and layout to the line-work down to how the gradients in the, in fashion, large logotype at the bottom of the footer are tied in by using texture. Was it in house, or an agency? I'd love to see some more of whoever's work it was
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#14Congrats on the launch! How do you guys compare with Datalab with regards to accuracy? https://www.datalab.to/
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#15Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#16Nice! I was already considering using reducto api. Will give this a try
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#17Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#18I saw your recent $24M series A and was kind of surprised to only see you launching now, congrats! YC seems to fund quite many document extraction companies, even within the same batch: - Pulse (YC W24): https://www.ycombinator.com/companies/pulse-3 - OmniAI (YC W24): https://www.ycombinator.com/companies/omniai - Extend (YC W23): https://www.ycombinator.com/companies/extend How do you differentiate from these? And h…
Generally speaking, my view on the space is that this was crowded well before LLMs. We've met a lot of the folks that worked on things like drivers for printers to print PDFs in the 1990s, IDP players from the last few decades, and more recent cloud offerings.
The context today is clearly very different than it was in the IDP era though (human process with semi-structured content -> LLMs are going to reason over most human data), and so is the solution space (VLMs are an incredible new tool to help address the problem).
Given that I don't think it's surprising that companies inside and outside of YC have pivoted into offering document processing APIs over the past year. Generally speaking we don't see differentiation in the sense of just feature set since that'll converge over time, and instead primarily focus on accuracy, reliability, and scalability, all 3 of which have a very substantive impact from last mile improvements. I think the best testament I have to that is that the customers we've onboarded are very technical, and as a result are very thorough when choosing the right solution for them. That includes a company wide roll out at one of the 4 biggest tech companies, one of the 3 biggest trading firms, and a big set of AI product teams like Harvey, Rogo, ScaleAI etc.
At the end of the day I don't see VLM improvements as antagonistic to what we're doing. We already use them a lot for things like an agentic OCR (correcting mistakes from our traditional CV pipeline). On some level our customers aren't just choosing us for PDF->markdown, they're onboarding with us because they want to spend more of their time on the things that are downstream from having accurate data, and I expect that there'll be room for us to make that even more true as models improve.
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#19Congrats on the launch guys, mobile website seems to be broken though.
Re: Launch HN: Reducto Studio (YC W24) – Build accurate document pipelines, fast
#20I saw your recent $24M series A and was kind of surprised to only see you launching now, congrats! YC seems to fund quite many document extraction companies, even within the same batch: - Pulse (YC W24): https://www.ycombinator.com/companies/pulse-3 - OmniAI (YC W24): https://www.ycombinator.com/companies/omniai - Extend (YC W23): https://www.ycombinator.com/companies/extend How do you differentiate from these? And h…
In this case, the Reducto team seems to have cloned us down to the small details [1][2], which is a bit disappointing to see. But imitation is the best form of flattery I suppose! We thought deeply about how to build an ergonomic configuration experience for recursive type definitions (which is deceptively complex), and concluded that a recursive spreadsheet-like experience would be the best form factor (which we shipped over a year ago).
> "How do you see the space evolving as LLMs commoditize PDF extraction?"
Having worked with a ton of startups & F500s, we've seen that there's still a large gap for businesses in going from raw OCR outputs —> document pipelines deployed in prod for mission-critical use cases. LLMs and VLMs aren't magic, and anyone who goes in expecting 100% automation is in for a surprise.
The prompt engineering / schema definition is only the start. You still need to build and label datasets, orchestrate pipelines (classify -> split -> extract), detect uncertainty and correct with human-in-the-loop, fine-tune, and a lot more. You can certainly get close to full automation over time, but it takes time and effort — and that's where we come in. Our goal is to give AI teams all of that tooling on day 1, so they hit accuracy quickly and focus on the complex downstream post-processing of that data.