Live data from Hacker News

Pdf2htmlEX – Convert PDF to HTML without losing text or format

github.com

41–50 of 51 posts

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#41

Can anyone recommend an equally good opposite (HTML to PDF)? wkhtmltopdf [0] is probably the most popular, but it's also ridiculously buggy. 0: https://code.google.com/p/wkhtmltopdf/

In OS X, you can print to PDF from every application.

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#43
post #28

I've actually been using this to convert large PDF files to HTML to be displayed in-browser. It's for my work, so I don't feel comfortable posting a link to the demo instance here. It is definitely the best solution I've found so far. The outputted HTML / CSS / images look almost identical to the source PDF. That being said, there are a few issues still: * One Gigantic (600kb) CSS file from a single PDF * Hundreds of…

Hey thanks for the info! 2nd & 3rd are in the future plan, as I'm still working on accuracy and speed. And #115( https://github.com/coolwanglu/pdf2htmlEX/issues/115 ) is about the 2nd issue. About the first one, I've not got an elegant solution yet, maybe a CSS file per page? Please file new issues at GitHub if you think it's necessary :)

I love this! Kudos for this awesome app.

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#44

Earlier quoted context omitted.

I heard that with careful optimization on the server side and a clever JS may solve this. So far the default UI just demostrates the ability of reading-while-downloading. The idea is that now the document becomes more controllable and accessible, say you can put Google Analytics in your resume written in LaTeX; or maybe an social reading service, where you can comment, annotate and share. Unlike PDF viewers, web brow…

Off topic, but is your username missing a "ke"?

um? why?

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#45

Interesting. So it converts all vector graphics to a background image per page, but keeps all text as browser-rendered on top of it. I guess I don't really see much practical purpose for it -- most browsers these days seem perfectly fine opening PDF files natively, after all. But it's a very cool technological demonstration. Maybe this could be some kind of bridge tool for generating sites with fancy typographical la…

It definitely has a few practical purposes. I have used this for a website for a small magazine. Their issue was that they didn't have resources to design for the web. This was a good solution, wherein they just needed to upload a PDF once an issue was out. And this provides a bit more flexibility from other PDF viewers - organize by articles, add social sharing, commenting per page/article etc. etc.

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#46

Earlier quoted context omitted.

Off topic, but is your username missing a "ke"?

um? why?

I believe he's wondering if your username is a reference to Cool Hand Luke.

http://www.imdb.com/title/tt0061512/

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#47
post #28

I've actually been using this to convert large PDF files to HTML to be displayed in-browser. It's for my work, so I don't feel comfortable posting a link to the demo instance here. It is definitely the best solution I've found so far. The outputted HTML / CSS / images look almost identical to the source PDF. That being said, there are a few issues still: * One Gigantic (600kb) CSS file from a single PDF * Hundreds of…

" * HTML semantics are non-existent

These are all relatively easy to fix, I believe. "

How? For example, how would you identify 's (or whatever this converter uses) to identify headers, and page headers/footers, or a ToC, or a preface? IMO this is an AI-hard problem, for which even the 'simple' approximation (statistics) is very hard due to the wide variety in inputs (a corpus trained for multi-column journal articles will most likely not work at all for books, although I haven't tried and would love to be proven wrong).

Use case: a working (i.e., preserving semantics) pdf-to-epub converter. This would, imho, be a killer product / service.

Re: Pdf2htmlEX – Convert PDF to HTML without losing text or format

#48
Question, does your public folder periodically delete files? I accidentally uploaded something confidential and it seems to be gone. I was wondering if this was a manual deletion or just expired since I still see files that were uploaded around the same time still there.
Post reply on HN