Skip to main content
cURL

Common Use Cases

  • Extract text from scanned PDFs and photos of paperwork while keeping table and column layout
  • Convert image-based invoices, contracts, or reports into readable text for search or data entry
  • Pull text from a specific region of a page, such as a totals box or letterhead
  • Turn scanned documents into text that keeps layout for analytics or AI prompts
Try it live: PDF to Text → API Tester — send a real request from your browser.

POST /v1/pdf/convert/to/text

Auto classification Of incoming documents: Use the Document Classifier endpoint to automatically sort/detect the class of the document based on keywords-based rules. For example, you can define rules to find which vendor provided the document to find which template to apply accordingly.

Request body

Attributes are case-sensitive and should be inside JSON for POST request. for example: { "url": "https://example.com/file1.pdf" } No query parameters accepted. Use the names as shown below, for example lineGrouping.
To see the request size limits, please refer to the Request Size Limits.
string
required
URL to the source file url attribute
string
The callback URL (or Webhook) used to receive the POST data. see Webhooks & Callbacks. This is only applicable when async is set to true.
string
HTTP auth user name if required to access source URL.
string
HTTP auth password if required to access source URL.
string
default:"all pages"
Specify page indices as comma-separated values or ranges to process (e.g. “0, 1, 2-” or “1, 2, 3-7”). The first-page index is 0. Use ”!” before a number for inverted page numbers (e.g. “!0” for the last page). If not specified, the default configuration processes all pages. The input must be in string format.
boolean
default:"false"
Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when lineGrouping is set to 1.
string
Defines coordinates for extraction. UsePDF Edit Add Helperto get or measure PDF coordinates. The format is {x} {y} {width} {height}.
string
default:"eng"
Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see Language Support. You can also use 2 languages simultaneously like this: eng+deu (any combination).
boolean
default:"false"
Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated.
string
Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: 1, 2, 3.
string
Password for the PDF file.
boolean
default:"false"
Set async to true for long processes to run in the background, API will then return a jobId which you can use with the Background Job Check endpoint. Also see Webhooks & Callbacks
string
File name for the generated output, the input must be in string format.
integer
default:"60"
Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from PDF.co Temporary Files Storage. The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using PDF.co Built-In Files Storage.
string
See Profiles for more information. This value is a JSON-encoded string.

Responses

Synchronous response
An asynchronous request returns a jobId and a reserved URL that should be used only after the job succeeds. Poll it via Background Job Check. Errors return error: true with a status code and message. See the response examples and Response Codes. Authentication and routing failures are separate: their response uses "status": "error" and provides the numeric HTTP code in errorCode (for example, 401).
Inconsistent URL Encoding in cURL Output: When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (&) may appear as \u0026 in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you’re parsing the response programmatically, your JSON parser will handle this conversion automatically.