Skip to main content
Inspiring
August 7, 2016
Question

Why does making my PDFs searchable compromise their visual quality?

  • August 7, 2016
  • 3 replies
  • 1025 views

I wrote some Internet pages to PDF. This results in non-searchable files (in contrast to saving Word documents to PDF, which maintains the OCR in the PDF).

Once written to PDF, I run OCR on the Internet pages that I saved to make them searchable. This significantly compromises the visual quality of the files. Is there a reason for this or a way to avoid it?

This topic has been closed for replies.

3 replies

jmt111Author
Inspiring
November 19, 2016

From within Google Chrome, I print the page to PDF using Adobe Acrobat X as the printer.

Actually, I just tried printing this page to PDF and the PDF file that was produced was searchable (i.e., I am able to search for words in the document and Acrobat finds them). When I originally posted, the file that was written to PDF lost its search-ability. I will post back if the issue recurs.

Inspiring
November 19, 2016

It is well known that "refrying" a PDF through the PDF Printer loses a lot of quality. You should not be doing this.

Refrying
PDFs
–

 the
good,
the
bad
and
the
ugly by Leonard Rosenthal

Legend
August 7, 2016

I think we need to talk about terms. There is no OCR when you go from Word to PDF!

OCR is looking at an image (usually but not always a scanned paper page), and using guesswork to figure out where the text is. Once Acrobat has the text it can keep the original image, replace it with the text, or do something smart in between.

Rather than trying to fix the OCR I'd look instead at why you need it. It's an error prone process, intended only when you have nothing but paper to work with. So, how do you make the web page? Do you use graphics for text? And how do you convert the web page to PDF? What version of Acrobat?