Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
51–60 of 66 posts
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#52I would like a tool that converts x months of credit card bills into a csv (the txn table from across PDFs and pages in each PDF) or something very easily.
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#53Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#54I would like a tool that converts x months of credit card bills into a csv (the txn table from across PDFs and pages in each PDF) or something very easily.
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#55Earlier quoted context omitted.
How do you guarantee that nothing in an extracted rent roll is hallucinated?
The same way you guarantee that a person manually typing the data never makes a mistake.
These tend to be easy to catch, even for the same person who's reviewing the data. They would see that rent steps looked strange (Y2 and Y4, but not Y3) or there was an order-of-magnitude difference in rent from one month to another.
AI can do something like invent reasonable-looking rent steps. They're designed to create output that seems reasonable, even if it's completely made up.
When humans are wrong, they tend to misread what's there, which is much less insidious than inventing something.
And if you have a human reviewer for all the work this AI does, what's the point of the AI in the first place? The human has become the source of truth either way.
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#56Tried the examples - they seem tailored for specific document types. I have two questions around that: (a) is their a "best-effort" extraction you can perform or plan to support if you don't know the document type? (b) do you plan to support extraction from academic papers, i.e., potentially multi-column, with images, tables that are either single column or span two columns, equations, etc.?
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#57Earlier quoted context omitted.
Execution is everything. Not to drop a link in someone else’s HN launch but I’m building https://therapy-forms.com and these guys are way ahead of me on UI, polish, and probably overall quality. I do think there’s plenty of slightly different niches here, but even if there were not, execution is everything. Heck it’s likely I’ll wind up as a midship customer, my spare time to fiddle with OCR models is desperately lim…
Just a heads up, but I tried to signup but the button doesn't seem to work.
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#58Congratulations on the launch! Its a crowded space but I think there is place for a good and accurate tool! Tried the examples - they seem tailored for specific document types. I have two questions around that: (a) is their a "best-effort" extraction you can perform or plan to support if you don't know the document type? (b) do you plan to support extraction from academic papers, i.e., potentially multi-column, with…
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#59I would like a tool that converts x months of credit card bills into a csv (the txn table from across PDFs and pages in each PDF) or something very easily.
Re: Launch HN: Midship (YC S24) – Turn PDFs, docs, and images into usable data
#60Can you speak to the accuracy, particularly of numerical value extraction, that you’re achieving? I have a use case for pulling tabular financial data out of PDFs and accuracy is our main concern with using AI for that type of task.