2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning
I want to try this tool with pages that Instapaper fails to grab.
11–20 of 77 posts
2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning
I want to try this tool with pages that Instapaper fails to grab.
I can't put down a person for effort, but this tool doesn't work at all. Over half the pages I tried are missing content entirely.
I tried it on this webpage : https://www.rt.com/news/357833-france-coca-cola-cocaine/ Not bad but the title of the article is missing. The tweet is missing too but I can't decide if it's a good or a bad thing. There should be options to remove pictures too I think. Honestly I'd be interested by a standalone product like this. I don't like the fact that you know everything I store.
The title of the article is the name of document, we understand that this is misleading so we can add it at the top of the document. For the Tweet, this is up for discussion for the moment our parser remove Media Card, we could add an option to enable or disable Media Card. For the picture this is an improvement we could do easily. Actually we do not know what document you generate, but I understand the need for end…
Later I tried a text export of the same page, some HTML remains ( elements).
Using the title of the page to name the zip would be nice too.
Interestingly, the webpage itself cannot be turned into a document. >We couldn't find any text for creating the document. Please send us the problematic link : https://documentcyborg.com/ via our contactus form. This was the first page I thought to try it on, so you might want to consider adding more text to your landing page so that it will work.
I can't put down a person for effort, but this tool doesn't work at all. Over half the pages I tried are missing content entirely.
Here's a freebie name that's (as of writing this) unregistered: page2doc.com
I can't put down a person for effort, but this tool doesn't work at all. Over half the pages I tried are missing content entirely.
It does seem to omit elements quite curiously. Just try using this submission as an example: https://news.ycombinator.com/item?id=12403661
1) the zipped download has always the same name, becomes quickly confusing if you're grabbing several pages. Maybe adding a timestamp or, better, a title snippet could help. 2) found html tags and incorrect whitespace when exporting to TXT this page: https://www.packtpub.com/packt/offers/free-learning I want to try this tool with pages that Instapaper fails to grab.
For the Instataper pages that fail if you could send us some domain name you wish : hn at documentcyborg.com and we will test the parser against it to make sure it works.
Earlier quoted context omitted.
The title of the article is the name of document, we understand that this is misleading so we can add it at the top of the document. For the Tweet, this is up for discussion for the moment our parser remove Media Card, we could add an option to enable or disable Media Card. For the picture this is an improvement we could do easily. Actually we do not know what document you generate, but I understand the need for end…
Sorry I forgot to mention : I generated a PDF. Later I tried a text export of the same page, some HTML remains ( elements). Using the title of the page to name the zip would be nice too.