Skip to main content
CURL
Try it live: PDF to JSON → API Tester — send a real request from your browser.

POST /v1/pdf/convert/to/json2

This endpoint can also be used with a specified profile to extract image data from a PDF into your JSON output.
You can extract hyperlinks from your PDF by using a profile to extract only link objects.

Request body

Attributes are case-sensitive and should be inside JSON for POST request. for example: { "url": "https://example.com/file1.pdf" } No query parameters accepted.
To see the request size limits, please refer to the Request Size Limits.
string
required
URL to the source file url attribute
string
HTTP auth user name if required to access source URL.
string
HTTP auth password if required to access source URL.
string
default:"all pages"
Specify page indices as comma-separated values or ranges to process (e.g. “0, 1, 2-” or “1, 2, 3-7”). The first-page index is 0. Use ”!” before a number for inverted page numbers (e.g. “!0” for the last page). If not specified, the default configuration processes all pages. The input must be in string format.
boolean
default:"false"
Unwrap lines into a single line within table cells in provided PDF documents. This is only applicable when linegrouping is set to 1.
string
Defines coordinates for extraction. UsePDF Edit Add Helperto get or measure PDF coordinates. The format is {x} {y} {width} {height}.
string
default:"eng"
Set the language for OCR (text from image) to use for scanned PDF, PNG, and JPG documents input when extracting text. see Language Support. You can also use 2 languages simultaneously like this: eng+deu (any combination).
boolean
default:"false"
Set to true to return results inside the response. Otherwise, the endpoint will return a URL to the output file generated.
string
Controls how lines of text are grouped when extracting data from a PDF. Line grouping within table cells. The available modes are: 1, 2, 3. For more information, see Line Grouping.
string
Password for the PDF file.
boolean
default:"false"
Set async to true for long processes to run in the background, API will then return a jobId which you can use with the Background Job Check endpoint. Also see Webhooks & Callbacks
string
File name for the generated output, the input must be in string format.
integer
default:"60"
Set the expiration time for the output link in minutes. After this specified duration, any generated output file(s) will be automatically deleted from PDF.co Temporary Files Storage. The maximum duration for link expiration varies based on your current subscription plan. To store permanent input files (e.g. re-usable images, pdf templates, documents) consider using PDF.co Built-In Files Storage.
string
Pass profile settings as a JSON-encoded string. See Profiles for more information.
You can use profiles to control the convert process and output of the JSON file.

Line Grouping Options

  • "1": GroupByRows – Each row is checked against the next row to see if they can be grouped together. Rows will only be grouped if all cells in the current row can be grouped with all cells in the next row. Useful when merging related content that spans multiple lines but belongs to the same logical row.
  • "2": GroupByColumns – Each cell is checked against the cell below it in the next row to determine if they can be grouped. Cells are grouped within the same column even if others can’t be grouped. Useful for columnar data where content in each column might span multiple lines.
  • "3": JoinOrphanedRows – Joins a row with a single cell to the previous row if there is no separator between them. Useful for handling cases with orphaned or misaligned content.

OCRImagePreprocessingFilters

To set image preprocessing filters, please use:

Responses

Synchronous response

HTTP response metadata

With inline: false (the default), the response contains a url to the generated JSON file, as shown above. With inline: true, the generated JSON content is returned in body instead. An asynchronous request returns a jobId and a reserved URL that should be used only after the job succeeds. Poll it via Background Job Check. Errors return error: true with a status code and message. See the response examples and Response Codes.

Generated JSON content

These fields belong to the generated JSON document. With inline: false, read them from the file at url. With inline: true, read them under the response’s body field. The paths below describe the default output structure; profiles can change the generated content.
Inconsistent URL Encoding in cURL Output: When using cURL to make API requests, the output JSON may show URL characters encoded as Unicode escape sequences. For example, the ampersand character (&) may appear as \u0026 in the cURL output. This is normal JSON encoding behavior and does not affect the validity of the URL. The URL will function correctly when used, as JSON parsers automatically decode these escape sequences. If you’re parsing the response programmatically, your JSON parser will handle this conversion automatically.