Skip to main content
Participant
January 29, 2025
Answered

Compressing a scanned file for size (not to recognize text) resulted in contents full of errors

  • January 29, 2025
  • 2 replies
  • 687 views

I used my Adobe Acrobat Pro account to compress a large scanned 98MB PDF file for size, not to recognize text, and the resulting compressed file is full or errors. It seems the fonts are all messed up with certain letters being replaced with white space (kind of like in the case of redacting), and so the contents are mostly ineligeble in this compressed file. It seems to me like it's reducing its size by literally deleting characters, which seems kind of absurd. What is going on that is causing this? I have not intended to or need to run text recognition on the file as I realize it is not the best quality text and there is handwritten notes throughout sporadically. How can I get a smaller size file without altering the contents (vs. reduction in clarity quality, which I am okay with!)?

 

Thanks!
-Yasmeen

Correct answer AnandSri

Hello,

 

I hope you're doing well, and we apologize for the delayed response and the trouble.

 

Is it specific to one file or is it with all the PDFs? When compressing scanned PDF files in Adobe Acrobat without performing Optical Character Recognition (OCR), it's essential to adjust specific settings to prevent content errors.

Use the 'Optimize Scanned PDF' Feature:

  • Open your scanned PDF in Adobe Acrobat.
  • Navigate to Menu/Tools > Optimize PDF.
  • In the toolbar, select Optimize Scanned PDF.
  • In the dialog box, ensure that Adaptive Compression is enabled. This setting helps reduce file size while maintaining content integrity. For detailed guidance, refer to this article.

Disable OCR During Optimization:

  • Within the same Optimize Scanned PDF dialog, look for the Recognize Text option.
  • Uncheck the Recognize Text option to prevent Acrobat from applying OCR during compression.

Check for Font Issues: Ensure the fonts used in the scanned document are correctly embedded. Sometimes, font issues can cause errors in the compressed file. Adobe provides guidance on how to manage fonts in PDFs: Fonts in PDFs. See this article for detailed information.

 

If you own the PDF and can access the source file, try recreating it using Acrobat, ensuring the fonts are correctly embedded.

 

Let us know how it goes.

Thanks,

Anand Sri.

2 replies

pdfilike
Participant
August 25, 2026

Hi there! What you are experiencing happens because Adobe Acrobat's internal PDF optimizer sometimes attempts to re-encode fonts or flatten hidden text layers during aggressive compression—even if OCR wasn't explicitly triggered. If the original PDF contains non-embedded fonts, custom symbols, or messy layers combined with handwritten notes, the compression engine can misinterpret characters, leading to those blank spaces or missing text blocks.

To get a smaller file size without altering or deleting the actual contents/text, while being completely okay with a slight drop in visual clarity, you can try these workarounds:

Rasterize the PDF (Convert to Images First): If the text integrity is getting corrupted by font re-encoding, the safest way to preserve what it looks like (including handwritten notes) is to convert the PDF pages into high-resolution images (PNG/JPG) and then re-compile them back into a clean PDF. Since it becomes a pure image-based document, Acrobat won't attempt any font substitution or text manipulation.

Adjust Acrobat Optimization Settings: If you are using the PDF Optimizer tool in Acrobat:

Go to Fonts in the settings menu and ensure you uncheck "Unembed any font" (this prevents the software from messing with your font structures).

Under Discard Objects and Discard User Data, be careful not to strip out crucial structural elements.

Alternative Compression Approach: If Acrobat continues to glitch on that specific 98MB file, you can try splitting the large document into smaller chunks or using alternative web-based tools that compress files strictly via image-stream downsampling rather than modifying internal font objects.

Hope this helps you safely reduce your file size without losing your text or notes!

Asad Ullah Khan
AnandSri
AnandSriCorrect answer
Legend
February 19, 2025

Hello,

 

I hope you're doing well, and we apologize for the delayed response and the trouble.

 

Is it specific to one file or is it with all the PDFs? When compressing scanned PDF files in Adobe Acrobat without performing Optical Character Recognition (OCR), it's essential to adjust specific settings to prevent content errors.

Use the 'Optimize Scanned PDF' Feature:

  • Open your scanned PDF in Adobe Acrobat.
  • Navigate to Menu/Tools > Optimize PDF.
  • In the toolbar, select Optimize Scanned PDF.
  • In the dialog box, ensure that Adaptive Compression is enabled. This setting helps reduce file size while maintaining content integrity. For detailed guidance, refer to this article.

Disable OCR During Optimization:

  • Within the same Optimize Scanned PDF dialog, look for the Recognize Text option.
  • Uncheck the Recognize Text option to prevent Acrobat from applying OCR during compression.

Check for Font Issues: Ensure the fonts used in the scanned document are correctly embedded. Sometimes, font issues can cause errors in the compressed file. Adobe provides guidance on how to manage fonts in PDFs: Fonts in PDFs. See this article for detailed information.

 

If you own the PDF and can access the source file, try recreating it using Acrobat, ensuring the fonts are correctly embedded.

 

Let us know how it goes.

Thanks,

Anand Sri.