Questions
Als letztes aktiv
Hi everyone,I’m using Adobe PDF Services API with NestJS to process scanned PDFs.My goal is to run OCR on the scanned file, then extract structured text and layout information for rendering in a Flutter frontend.However, I’ve noticed that sometimes the entire page is treated as one large image, and no text is extracted — even though the PDF clearly contains readable text after OCR.This causes layout issues in my Flutter app, where the image layer overlaps or replaces text.Implementation details const pollingURL = await pdfServices.submit({ job }); const pdfServicesResponse = await pdfServices.getJobResult({ pollingURL, resultType: OCRResult });Extract step:const params = new ExtractPDFParams({ elementsToExtract: [ExtractElementType.TEXT, ExtractElementType.TABLES], addCharInfo: true, getStylingInfo: true, elementsToExtractRenditions: [ ExtractRenditionsElementType.FIGURES, ExtractRenditionsElementType.TABLES, ], });! The issueFor some scanned PDFs, Adobe
I’m facing an issue reading text from the PDF using the Adobe's OCR API. Even after performing preprocessing through the API, the text remains unreadable. Could anyone please suggest how to extract or read text from P&ID symbols effectively?
I have a side project that allows users to do question-answering over their collection of pdfs. I'm using adove pdf APIs to extract the content of the pdf. I'm using adobe because it gives good structured output.Today I opened the billing page on my aws and seen that i have a bill of $3,400. I have 3,000 registered users on my tool and they may have uploaded 20,000 pdfs combined. So even in that case $3,400 bill for 20,000 pdf processing is too much.I seen the pricing of the service and it is somewhere around $1 per 100 pages. For other services adobe charge per 50 pages but for extract api it charges per 5 pages. Its not good. Adobe should lower their costs or charge as per 50 pages $0.05 just like their other APIs.
Is it possible to give the upload temp file url as input to the compression API. Getting a 500 error for that. Below is my code:- $ch = curl_init('https://pdf-services-ue1.adobe.io/operation/compresspdf');curl_setopt($ch, CURLOPT_POST, true);curl_setopt($ch, CURLOPT_HTTPHEADER, ['x-api-key: ' . $client_id,'Authorization: Bearer ' . $accessToken,]);curl_setopt($ch, CURLOPT_POSTFIELDS, ['file' => new CURLFile($upload['tmp_name'], 'application/pdf', $upload['name'])//temp file]);curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);curl_setopt($ch, CURLOPT_HEADER, true); // capture headerscurl_setopt($ch, CURLOPT_VERBOSE, true); // (optional) for debugging$response = curl_exec($ch);$header_size = curl_getinfo($ch, CURLINFO_HEADER_SIZE);$header = substr($response, 0, $header_size);$body = substr($response, $header_size);curl_close($ch);
Hello, Sorry if this is not in the right forum to voice this issue in, but to start off, I have been running a Power Automate Cloud flow using both the "Convert PDF to List of Images" and the "Convert PDF to Image" actions from the Adobe PDF Services Connector to turn a PDF into a PNG file.I've noticed that starting about 3 weeks ago the resulting PNG quality from the conversions have fallen. Fallen, as in the saturation level has been reduced in the PNG leaving colors looking darker, and the text is more pixelated in the PNG than in the original PDF. Has this happened to anyone else too, and has there been any fixes to this yet?
Hello everyone,i hope this is the right place to ask my specific question.As questioned long time ago on Community Board: https://community.adobe.com/t5/acrobat-services-api-discussions/math-equation-extraction-from-pdf/td-p/13033193 https://community.adobe.com/t5/acrobat-services-api-discussions/pdf-extract-api-math-image-extraction/td-p/14512088 the question is still ongoing. How to extract math formulars as separate image from a PDF.As example we can take Nr. 2 of the preceding question.In the Example Project i don't found my anwer: https://github.com/adobe/PDFServices.NET.SDK.Samples Perhaps there other solutions, or it is not now implemented as I expected ...Could anyone help me?
Hi Team,I’d like to understand how the Adobe Acrobat PDF Extract API identifies the structure of a PDF document.Does it rely on OCR-based detection, or does it use other layout or heuristic-based methods for identifying elements such as paragraphs, lists, tables, and headings?Could you please share some insights or documentation references about how the API determines and classifies these structural elements?Thanks,Sathish
Hi everyone, I client of mine wants to electronically seal PDFs in SharePoint.Their preference would be to use Power Automate with the Adobe PDF Services Connector and SwissSign as the TSP. According to the AATL List, SwissSign is supported (Adobe Approved Trust List Members, Acrobat) However, when using the PowerAutomate action to electronically seal a document, the TSP-Name dropdown only allows one of these choices:- ENTRUST- INTESI GROUP- GLOBAL SIGN- TRUST PRO How can we use SwissSign as the TSP with the PowerAutomate action? Thanks for any pointers!
Hi, I’ve been testing the PDF Services API (exportpdf operation) to convert PDFs into PowerPoint (PPTX), and I’ve run into an issue with fonts.In our PDFs, we’re using Roboto and Inter-Bold, and I’ve double-checked that they’re embedded in the PDF metadata. When I send the file through the API (tested with both the Node.js SDK pdfservices-node-sdk and the Python SDK, hitting pdf-services.adobe.io/operation/exportpdf), the output PPTX doesn’t keep those fonts. Instead, every text element defaults to Arial, and the spacing/layout looks off as well.I went through the API documentation and options but couldn’t find anything about controlling or overriding fonts. From what I can see, there aren’t any parameters exposed to handle this.Has anyone here run into the same problem? Is font preservation supported in PDF → PPTX conversions, or is it expected that custom fonts get replaced? Any guidance or workarounds would be super helpful — we need to preserve brand typography in client prese
The service has been working great most of the time but for this one attached PDF we cant process it. After about 90 seconds of polling every 3 seconds for the job status we get this: {"error"=>{"code"=>"REQUEST_TIMEOUT", "message"=>"The operation has timed out, please try after some time.;", "status"=>500}, "status"=>"failed"} Unfortunately this message doesn't tell us anything about what might be wrong with the extract or what to do. I also checked the console logs and it doesnt appear to actually log any of the PDF requests we have made, which is odd. I would have thought there would be a log of API requests with details. Anyway, anyone know why this PDF is causing issues?
We are trying out ”Adobe PDF services API”.Want to convert PDF to Docx format.We are evaluating ”Adobe PDF services API” for PDF to Docx format.We ran into some bugs, kindly help here to solve this:1. The table of contents are not navigable2. The Disclaimer icon in the first page is different3. The numbered tiles comes without number (for sub-titles like 3.1 / 3.2 etc)The source PDF and the corresponding docx are attached here.PDF --> TaggedPDF.pdf. (we tried tagging feature also)Docx --> docx-from-tagged-pdf.docx. (docx file generated after tagging.)Could you kindly help here to solve this.
Hi ,Whenever I am trying to create dynamic html to pdf. I am getting bad address error.Earliar it was working fine but now I am getting issue.Whenever I am doing download polling to asset I am getting bad address Error.{'error': {'code': 'BAD_ADDRESS', 'message': 'Bad Address, request terminated; requestId=xxxxxxxxxxxxxxxx', 'status': 400}, 'status': 'failed'}
Hi everyone, I'm trying to test the following "PDF to Markdown" operation. Specifically, I'm using the API as follows: curl --location 'https://pdf-services-ue1.adobe.io/operation/pdftomarkdown' \ --header 'X-API-Key: {KEY}' \ --header 'Authorization: Bearer {TOKEN}' \ --header 'Content-Type: application/json' \ --data '{ "assetID": "urn:aaid:AS:UE1:88b99474-afee-4d97-97e3-47de96ed6c06", "getfigures": true }'But I always get a 400 Band Request error with the following content:{ "error": { "code": "INVALID_REQUEST_FORMAT", "message": "Invalid request format." } }
Hi, I would like to extract English and Japanese text from tables in PDF, which is made by scanning printed paper. Original PDF:When I use Acrobat Pro DC(Convert to xlsx), the output quality is good. But when I use Adobe PDF Services API(extract_text_table_info_figures_tables_renditions_from_pdf.py), I get garbled text for the Japanesse text. I know the API is currently optimized for English language content. But is it possible to improve this API to the quality of Acrobat Pro DC?Beacuse I need to convert many PDFs, so I need CLI solution.I attach the original Excel and PDF file, so please use these as test data.Thank you.
How to Rotate a pdf while viewing it in PDF Embed APILike Clockwise, AntiClockwise or 180 degrees rotate. sample example is given below:
Hi,I'm using CreatePDFJob , to convert my pptx to pdf. My submitted job goes in hung state, and never responds.I have attached a pptx, which has only one slide, which is causing the issue. The slide has image, which I suspect is root cause.Please let me know how solve this issue.thanks
Hi there! I'm trying to get the Embed API set up on our site, but I'm running into some issues. It seems like it works the first time you open the page, but on reload, it disappears sometimes. I have it set up here: https://annescollege.fsu.edu/embed-testWhen I'm logged in to edit the site, it won't show, regardless of how many times I reload the page. We're using Drupal to manage the site. Also, when I'm viewing on an iPhone, it won't go full screen on both Chrome and Safari. Attaching screenshots of this issue. I'd appreciate any insight, thanks!
Hi, I'm currently studying on a project which I need to insert a single charge invoice PDF to the last page of some specific invoice PDF through Power Automate. The source invoice PDF file includes many invoices which has different page counts. So what I plan to do is to split the whole PDF file into invidivual PDF. I have do some research on web and find a similar example. https://medium.com/adobetech/split-pdfs-based-on-content-with-adobe-pdf-extract-service-with-microsoft-power-automate-a08dc5fbafaa I try to follow the same workflow and change the searchArray logic. In our invoice, there is "Page 1 of X" will be shown on each page, so I try to use this to identify the page count. However, when the flow run to the SearchArray, it got the below error. Action 'searchArray' failed: The execution of template action 'searchArray' failed: The evaluation of 'query' action 'where' expression '@startsWith(item()?['Text'],'PAGE 1 OF')' failed: 'The t
Hello Adobe Team,I’m working with scanned invoices and using two APIs together:OCR API → to make the scanned PDF editable/searchable.Extract API → with parameters: const params = new ExtractPDFParams({ elementsToExtract: [ExtractElementType.TEXT, ExtractElementType.TABLES], addCharInfo: true }); This works, but I’ve noticed unexpected results when reviewing the JSON and trying to re-render the PDF:The JSON output includes BBox attributes that add rectangular boxes around text and table elements.When rendering from this JSON in Flutter, extra borders appear that do not exist in the original scanned PDF (e.g. double borders around tables, boxes around text).It seems the API is treating every detected line or text area as a bounding rectangle, not just the actual drawn table/line borders from the original file.Example: a single drawn line in the PDF becomes a rectangle in the JSON.This makes it impossible to distinguish between real visual borders vs. bounding boxes used for OCR posi
I’ve been using a custom implementation to toggle the Adobe comment panel by setting theshowCommentsPanel value. This approach has always worked reliably, but it has suddenly stopped functioning as expected.Could you please confirm if there were any recent changes that may have affected this behavior? Additionally, what is the recommended way to programmatically toggle the comment panel now?Thank you for your guidance.
HI, i'm working on this AI agent tool called aiagent.surf and i am wondering can i change the output format of the pdf api into markdown format?The output format of acrobat pdf api is not suitable for LLM applications like AI agent. I tried converting the output programmatically but this task is so complex that i gave up.It would be cool if pdf api can combine seperate excel files and image files into one single markdown file for each pdf page.If there should be an inbuilt option to get output as markdown then it would be greatly helpful for ai chat pdf applications.There is service from Mathpix that does it but i find acrobat pdf api to be better in OCR.Any advice?
Hi We have been getting no PDF processing of .xls files for the last week - spotted 02/September/2025.I can generate a token, assigned an upload URI, receive a location, and upload the xls successfully.Monitoring the process returns "in progress" until default time-out of 10min.So, the file goes up but isn't coming down.I can only assume the issue lies on Adobe's side of the fence. I have tried the following endpoints:https://pdf-services.adobe.io https://pdf-services-ue1.adobe.io https://pdf-services-ew1.adobe.io We are on Free Tier and have used only 51/500 transactions this month.I notice that the storage points are hosted by AWS. Is this a factor?I check the Status dashboard constantly but always shows everthing is ok.Is anyone else having this problem? RegardsAlistair
Can\'t connect to HTTPS URL because the SSL module is not available.". This is the erros I got when I try to run python src/extractpdf/extract_txt_from_pdf.py on pycharm
Hi Community members,I am exploring the pdf OCR and EXTRACT APIs ( OCR, EXTRACT )I have a Scanned pdf so to make it editable i applied the OCR and then for the pdf style and content information i am using the Extract API ( Extract Text and Tables and Character Bounding Boxes (w/ Renditions) ) I have used this api into node like below const params = new ExtractPDFParams({ elementsToExtract: [ExtractElementType.TEXT, ExtractElementType.TABLES], addCharInfo: true });But the JSON which is extracted contains some extra info like added some attributes (boxes into the elements) but if you look into the original pdf then there are no boxes then why those added?
Trying to extract health informtion using PDF Services API and facing error 1) Extraction error coming from the pdf file uploaded (in this case a health report), 2) Create asset arror. Please support wiht your insights, I need to ship this feature urgently
Remix with Firefly Community Gallery
Thousands of free creations to fall in love with and remix in Firefly.
Sie haben bereits einen Account? Anmelden
Noch kein Konto? Konto erstellen
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.