Hello @corey_6791
Thanks for the detailed information.
Make sure the PDF has a real text layer and a sane structure
-
Run OCR (if text isn’t selectable)
In Acrobat: All tools> Scan & OCR> Recognize Text (pick language, pages). Re‑export to TXT/HTML afterwards. This prevents blank exports. See this article for more information: https://adobe.ly/48aQoMh
- Auto‑tag for structure
In Acrobat: All tools>Accessibility> Autotag Document. Good tags help downstream converters (including Calibre) interpret reading order, headings, lists, and tables. See this article for more information: https://adobe.ly/4an7EPK
- Use “Reflow” to verify reading order
Menu > View > Zoom> Reflow. If Reflow appears incorrect, correct the order via Accessibility> Reading Order before exporting.
Export Accessible HTML (single page) or simplified DOCX (flowing text, no fixed layout) >> Convert in Calibre with heuristics on.
I hope this helps.
Thanks,
Anand Sri.