It is a know and already discussed on Drupal.org fact that PDF documents of image type (without text layers) are not searchable. Would be awesome if this module could process files through OCR during upload. Google already uses OCR to search PDF files (http://googleblog.blogspot.com/2008/10/picture-of-thousand-words.html, http://googlesystem.blogspot.com/2008/10/google-uses-ocr-to-index-pdf-fi...) Why this can't be implemented on Drupal?

Comments

yngens’s picture

Issue summary: View changes

Change the text to concentrate on the subject.

amontero’s picture

Issue summary: View changes

Recent Tika versions do OCR for images, IIRC.