Live data from Hacker News

Show HN: Copyfish – Extract text from images, videos or PDF

addons.mozilla.org

31–40 of 93 posts

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#31
post #17
post #14

Earlier quoted context omitted.

Well, at least this confirms that the screenshots are not manipulated ;) The tricky part for the OCR in this example is the diverse background, as the Chinese characters are directly inside the movie. Your comment is interesting, as the original motivation for creating the Copyfish extension was to help me watch Chinese movies. So I can confirm that for this purpose, it works fine. Of course, once in a while it gets…

> as the Chinese characters are directly inside the movie. Yep, same with TV shows, and soft-copies of transcripts are difficult to come by, hence my interest in something like this. I just watched the video. When used on a video does it keep a history of all OCRed text? Finally, you might also like to try posting this on http://www.chinese-forums.com If it mostly works well for TV and films, I'm sure there will be q…

> When used on a video does it keep a history of all OCRed text?

Not yet - but this feature is already on my todo list ;)

Thanks for the hint about the chinese forums!

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#33
post #31
post #17

Earlier quoted context omitted.

> as the Chinese characters are directly inside the movie. Yep, same with TV shows, and soft-copies of transcripts are difficult to come by, hence my interest in something like this. I just watched the video. When used on a video does it keep a history of all OCRed text? Finally, you might also like to try posting this on http://www.chinese-forums.com If it mostly works well for TV and films, I'm sure there will be q…

> When used on a video does it keep a history of all OCRed text? Not yet - but this feature is already on my todo list ;) Thanks for the hint about the chinese forums!

> Not yet - but this feature is already on my todo list ;)

Another interesting feature would be to do some sort of statistical analysis of Chinese text being OCRed and then combining that with possible characters suggested by the OCR. This would almost certainly prevent the mistake in the last two characters of the Chinese movie screenshot.

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#34

Give me an api end point to send an image to, and a text response. Ill hand you cash.

There are a ton of these now. Google provides OCR as part of their machine vision API. AWS has similar with Rekognition. As others have mentioned, there are dozens of others on less well known platforms.

Rekognition from Amazon doesnt have OCR as far as I remember

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#36
post #32
post #25

Could you add an email address to your HN profile so I can contact you?

Done. In addition, the email listed on https://github.com/A9T9 also reaches me.

> Done. In addition, the email listed on https://github.com/A9T9 also reaches me.

Neat! Brother. +1 =100 Ace

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#37
post #11

Is the OCR-extraction performed in the client? if its transferred to a server then people should be aware of this so sensitive data from documents/pdf is not submitted.

Yeah apparently it uses https://ocr.space/ , deal-breaker for me.

Why is it a deal breaker?

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#38

Give me an api end point to send an image to, and a text response. Ill hand you cash.

There are a ton of these now. Google provides OCR as part of their machine vision API. AWS has similar with Rekognition. As others have mentioned, there are dozens of others on less well known platforms.

Actually, based on my tests, there are only a few good services:

Abbyy (best recognition rate but by far most expensive), Google Cloud Vision (second best recognition rate), Microsoft OCR and... our OCR.space service with a very generous free tier and a competitive priced PRO tier.

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#39
post #11

Earlier quoted context omitted.

Yeah apparently it uses https://ocr.space/ , deal-breaker for me.

Why is it a deal breaker?

It's a deal breaker because THAT'S NONE OF YOUR DAMN BUSINESS, and that also goes for Copyfish. It smells fishy to me, and _promises_ never kept prying eyes away secret documents. People who handle confidential documents should never use SaaS. It's an issue of trust, and Copyfish deserves none.

Re: Show HN: Copyfish – Extract text from images, videos or PDF

#40
We need to evolve a grammar for describing privacy implications, because proper classification of this software would allow it to be marked as malware/spyware.

It is beyond irresponsible for mozilla to do nothing to prevent this malware from being recommended on their platform.

Post reply on HN