> This page is for version v1 (default).
> For other versions, use one of these documentation indexes:
> - v1 (default): https://docs.plextera.com/v-1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.plextera.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.plextera.com/_mcp/server.

# Document Extraction

Document Insights extracts structured field values from a document. The flow is asynchronous: you submit a document, receive an extraction ID, then read the result once processing finishes.

A typical integration looks like this:

1. Submit a document.
2. Poll or subscribe for the terminal state.
3. Read the output.
4. Optionally submit feedback.

All requests require an API key in the `Authorization` header. See [Authentication](/guides/introduction/authentication) for how to obtain and send one.

## Configuration, routing, and output

An extraction request selects a source document.
It does not define the fields to extract.
Your workspace's Document Insights configuration determines:

* which documents it accepts;
* how a document is routed;
* which fields, types, and nested structures appear in `output.fields`.

Before integrating, review the relevant Document Insights configuration in Plextera to identify the expected output schema and routing mode.
If you do not know how to access it, ask your workspace administrator or Plextera.

If default routing is configured, omit `docType`.
Other custom labels remain valid and are still returned for correlation.
If `docType` routing is configured, use the exact value from the Document Insights configuration in Plextera.
If you do not know where to find it, ask your workspace administrator or Plextera.
The v1 API does not currently expose a discovery endpoint for available configurations or `docType` values.

When no configuration matches, the extraction reaches `FAILED` with `error.code` set to `DOCUMENT_NOT_ROUTED`.
See [Handle failures](#handle-failures) for the corrective action.

## Submit a document

Choose one of two submission styles, depending on whether you want to manage the file separately.

The examples below use default routing. If your workspace requires `docType`, add it as shown in [Labels and docType](#labels-and-doctype).

### Option A: Extract from a file or URL

Use this when the document already lives in Plextera File Service or is reachable at an HTTPS URL. This is also the right choice when you want a reusable `fileId` you can extract from more than once.

#### Stored file

To extract from a stored file, first upload it with `POST /files`, then reference the returned `id`.

```bash
# 1. Upload the file
curl -X POST https://api.plextera.com/api/public/v1/files \
  -H "Authorization: api-key YOUR_API_KEY" \
  -F "file=@invoice.pdf"
```

```json
{
  "id": "file_01JY7M4ZVX5R1P3M3Q0TA1S7ZM",
  "fileName": "invoice.pdf",
  "mimeType": "application/pdf",
  "size": 63877
}
```

```bash
# 2. Start the extraction
curl -X POST https://api.plextera.com/api/public/v1/document-insights/extractions \
  -H "Authorization: api-key YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "document": {
      "fileId": "file_01JY7M4ZVX5R1P3M3Q0TA1S7ZM"
    }
  }'
```

#### Document URL

To extract from a URL, pass `document.url` with `document.fileName`. No prior upload is needed; Plextera downloads the document.

```bash
curl -X POST https://api.plextera.com/api/public/v1/document-insights/extractions \
  -H "Authorization: api-key YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "document": {
      "url": "https://example.com/documents/invoice.pdf",
      "fileName": "invoice.pdf"
    }
  }'
```

### Option B: Upload and extract in one request

Use this when you do not need a reusable `fileId` and want a single round trip. Send the file as `multipart/form-data` to `POST /document-insights/extractions/upload`.

```bash
curl -X POST https://api.plextera.com/api/public/v1/document-insights/extractions/upload \
  -H "Authorization: api-key YOUR_API_KEY" \
  -F "file=@invoice.pdf"
```

The `file` part is required. To add routing or correlation labels, send the optional `labels` part as a JSON object as shown below.

## What you get back

Every submission method returns the same extraction object, accepted asynchronously with status `QUEUED`.

```json
{
  "id": "69654f0bc073ef404baec649",
  "status": "QUEUED",
  "outputAvailable": false,
  "createdAt": "2026-04-07T10:05:04Z",
  "updatedAt": "2026-04-07T10:05:04Z",
  "document": {
    "fileId": "file_01JY7M4ZVX5R1P3M3Q0TA1S7ZM",
    "fileName": "invoice.pdf"
  }
}
```

> **Tip**
>
> Store the `id`. This is the `extractionId` you use to poll for results, submit feedback, and correlate events.

All timestamps in extraction responses are UTC ISO 8601 strings, for example `2026-04-07T10:22:00Z`.

## Labels and `docType`

Both submission methods accept optional `labels`.
Custom labels are string key-value pairs that you define for correlation, such as a case ID or source system.
Plextera returns them on every response and event for the extraction.
Custom labels do not change which fields are extracted.

For a JSON extraction request:

```json
{
  "document": {
    "fileId": "file_01JY7M4ZVX5R1P3M3Q0TA1S7ZM"
  },
  "labels": {
    "customerDocumentId": "doc-42",
    "sourceSystem": "erp"
  }
}
```

`docType` is different: it is a reserved label that drives document routing.
Do not invent this value.
Use the exact value from the relevant Document Insights configuration in Plextera.
If you do not know where to find it, ask your workspace administrator or Plextera.
When your workspace uses default routing, omit `docType`; you can still send any custom labels you need for correlation.

```json
{
  "labels": {
    "customerDocumentId": "doc-42",
    "docType": "YOUR_CONFIGURED_DOCUMENT_TYPE"
  }
}
```

For a multipart upload, encode the same object as the JSON `labels` part:

```bash
curl -X POST https://api.plextera.com/api/public/v1/document-insights/extractions/upload \
  -H "Authorization: api-key YOUR_API_KEY" \
  -F "file=@invoice.pdf" \
  -F 'labels={"docType":"YOUR_CONFIGURED_DOCUMENT_TYPE","customerDocumentId":"doc-42"};type=application/json'
```

You can attach up to 50 labels. Keys are limited to 64 characters and values to 512 characters; keys and values must be non-empty after trimming.

## Get the result

An extraction moves through one of these lifecycles:

```mermaid
stateDiagram-v2
    direction LR
    [*] --> QUEUED
    QUEUED --> PROCESSING
    QUEUED --> REJECTED: document rejected
    QUEUED --> FAILED: routing failed
    PROCESSING --> COMPLETED
    PROCESSING --> FAILED: processing or data validation failed
    PROCESSING --> REJECTED: document rejected
    COMPLETED --> [*]
    FAILED --> [*]
    REJECTED --> [*]
```

`COMPLETED`, `FAILED`, and `REJECTED` are terminal. You can receive the result by polling or by subscribing to events.

### Poll for output

Call `GET /document-insights/extractions/{extractionId}` until the status is terminal.
A reasonable cadence is every 2-5 seconds for the first minute, then every 15-30 seconds.
Most documents finish within a couple of minutes; large or multi-page documents can take longer.

```bash
curl https://api.plextera.com/api/public/v1/document-insights/extractions/69654f0bc073ef404baec649 \
  -H "Authorization: api-key YOUR_API_KEY"
```

When `status` is `COMPLETED`, `outputAvailable` is `true` and the response carries the `output` object. When `status` is `FAILED` or `REJECTED`, it carries an `error` object instead.

## Read the output

A completed extraction returns `output.fieldCount` and an `output.fields` array. Each field carries its extracted `value` plus extraction details such as confidence, page, and placement.

```json
{
  "id": "69654f0bc073ef404baec649",
  "status": "COMPLETED",
  "outputAvailable": true,
  "createdAt": "2026-04-07T10:05:04Z",
  "updatedAt": "2026-04-07T10:22:00Z",
  "completedAt": "2026-04-07T10:22:00Z",
  "document": {
    "fileId": "file_01JY7M4ZVX5R1P3M3Q0TA1S7ZM",
    "fileName": "invoice.pdf",
    "mimeType": "application/pdf",
    "size": 63877,
    "pageCount": 1,
    "contentUrl": "https://files.plextera.com/acme/2026/04/07/invoice_9f1c8b2e.pdf?Expires=1775900000&KeyName=cdn-key&Signature=Zm9vYmFyc2lnbmF0dXJl"
  },
  "output": {
    "fieldCount": 3,
    "fields": [
      {
        "id": "field_01",
        "name": "vendorName",
        "type": "text",
        "value": "Acme Industries Ltd",
        "metadata": {
          "extracted": true,
          "confidence": 1.0,
          "page": 1,
          "placement": { "x": 43, "y": 12, "width": 22, "height": 4 }
        }
      },
      {
        "id": "field_02",
        "name": "invoiceNumber",
        "type": "text",
        "value": "INV-0042",
        "metadata": {
          "extracted": true,
          "confidence": 0.98,
          "page": 1,
          "placement": { "x": 316, "y": 48, "width": 147, "height": 10 }
        }
      },
      {
        "id": "field_03",
        "name": "totalAmount",
        "type": "text",
        "value": "1,024.33 USD",
        "metadata": {
          "extracted": true,
          "confidence": 0.99,
          "page": 1,
          "placement": { "x": 316, "y": 209, "width": 147, "height": 10 }
        }
      }
    ]
  }
}
```

Each field's `id` is what you pass as `fieldId` when submitting feedback. For the full field schema, see the [API reference](/api/api-reference/document-insights).

### Download the source document

The `document` object carries file metadata, including a standard MIME value such as `application/pdf`, plus `contentUrl` - a short-lived pre-signed link to download the original document straight from storage. The MIME value is omitted when Document Insights cannot determine it. The download link is minted on each Get extraction request and expires, so fetch it promptly and don't persist it; request the extraction again to get a fresh one. Extraction events carry the same `contentUrl`, minted when the event is created. It is not returned by [List extractions](#list-extractions), and is omitted when the link cannot be resolved. The [Files guide](/guides/core-guides/files) shows the same download pattern for files fetched by `fileId`.

## Delete an extraction

Delete an extraction you no longer need with `DELETE /document-insights/extractions/{extractionId}`.

```bash
curl -X DELETE https://api.plextera.com/api/public/v1/document-insights/extractions/69654f0bc073ef404baec649 \
  -H "Authorization: api-key YOUR_API_KEY"
```

A successful delete returns `204 No Content`. Only terminal extractions can be deleted - deleting one that is still `QUEUED` or `PROCESSING` returns `409 CONFLICT`; wait for a terminal status first. Deletion is permanent: the extraction disappears from `GET` and list responses and from Plextera review tooling.

## Handle failures

`FAILED` and `REJECTED` both return an `error` object with a machine-readable `code` and a human-readable `message`. The distinction matters when deciding whether to retry.

| Status     | What it means                                                                                                                                                | Retry guidance                                                                                           |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------- |
| `REJECTED` | The document could not be accepted. Causes include duplicate content, a broken, empty, or oversized file, an unsupported format, or an unsupported language. | Read `error.message` and correct the document before retrying.                                           |
| `FAILED`   | Routing, processing, or extracted-data validation could not complete.                                                                                        | Use `error.code` to distinguish a configuration problem from a potentially temporary processing failure. |

Common terminal error codes:

| Code                         | Status     | What to do                                                                                                                                                                           |
| ---------------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `DOCUMENT_NOT_ROUTED`        | `FAILED`   | Verify the exact `labels.docType` value. If you omitted `docType`, ask your workspace administrator whether default routing exists. Retrying the unchanged request will not help.    |
| `DOCUMENT_REJECTED`          | `REJECTED` | Read `error.message` for the concrete reason. For a duplicate, reuse the existing extraction or submit different content. For an invalid document, correct the file before retrying. |
| `DOCUMENT_EXTRACTION_FAILED` | `FAILED`   | Retry once with backoff if the failure may be transient. If it repeats, contact Plextera support.                                                                                    |

#### Duplicate

```json
{
  "id": "69654f0bc073ef404baec649",
  "status": "REJECTED",
  "outputAvailable": false,
  "completedAt": "2026-04-07T10:05:20Z",
  "error": {
    "code": "DOCUMENT_REJECTED",
    "message": "Duplicate document. Original document ID: 69654f0bc073ef404baec600"
  }
}
```

#### Not routed

```json
{
  "id": "69654f0bc073ef404baec649",
  "status": "FAILED",
  "outputAvailable": false,
  "completedAt": "2026-04-07T10:05:20Z",
  "error": {
    "code": "DOCUMENT_NOT_ROUTED",
    "message": "No Document Insights configuration matched this document. Verify labels.docType or contact your workspace administrator."
  }
}
```

#### Failed

```json
{
  "id": "69654f0bc073ef404baec649",
  "status": "FAILED",
  "outputAvailable": false,
  "completedAt": "2026-04-07T10:18:00Z",
  "error": {
    "code": "DOCUMENT_EXTRACTION_FAILED",
    "message": "The document could not be processed."
  }
}
```

When contacting support, include the extraction `id` plus `error.code` and `error.message`.
These values give support enough context to locate the extraction and diagnose the reason.

## Receive events instead of polling

If your application exposes a webhook endpoint, subscribe to extraction events and let Plextera push the terminal result to you. The completed event payload contains the same extraction model returned by `GET /document-insights/extractions/{extractionId}`, including `output`.

Recommended event types:

* `document-insights.extraction.completed` includes `output`.
* `document-insights.extraction.failed` includes `error`.
* `document-insights.extraction.rejected` includes `error`.

See [Event Subscriptions](/guides/core-guides/event-subscriptions) for endpoint setup, signature verification, and retry behavior.

## List extractions

Use `GET /document-insights/extractions` to review history or build a status view. It returns a paged summary of each extraction without the full `output`, and the `document` summary here does **not** include the `contentUrl` download link. Fetch a single extraction by ID when you need the field values or a download link.

```bash
curl "https://api.plextera.com/api/public/v1/document-insights/extractions?status=COMPLETED&from=2026-04-01T00:00:00Z&to=2026-04-30T23:59:59Z&sortBy=createdAt&sort=desc" \
  -H "Authorization: api-key YOUR_API_KEY"
```

Filter by `status`, narrow by a `from` / `to` created-time window, and order with `sortBy` and `sort`. Results are paged with zero-based `page` and `size`, and every response includes a `pageInfo` object. See the API reference for the complete parameter list and defaults.

### Filter by labels

Filter on the labels you submitted with each extraction using the `labels[key]=value` parameter. Use `labels[key]=` (empty value) to match extractions that have the label, or `labels[key]=value` for an exact match. Multiple keys are combined with AND.

```bash
curl --get https://api.plextera.com/api/public/v1/document-insights/extractions \
  -H "Authorization: api-key YOUR_API_KEY" \
  --data-urlencode "labels[docType]=YOUR_CONFIGURED_DOCUMENT_TYPE" \
  --data-urlencode "labels[customerDocumentId]=doc-42"
```

This returns extractions with your configured `docType` **and** `customerDocumentId=doc-42`. Label filters combine with the other filters above.

## Submit feedback

When an extracted value is wrong or needs review, submit feedback against the extraction. Feedback is surfaced to the team that configures your workspace and is used to improve extraction quality.

```bash
curl -X POST https://api.plextera.com/api/public/v1/document-insights/extractions/69654f0bc073ef404baec649/feedback \
  -H "Authorization: api-key YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "type": "ERROR",
    "fieldId": "field_01",
    "message": "Vendor name should be Acme Industries Ltd."
  }'
```

| Field     | Description                                                                                                  |
| --------- | ------------------------------------------------------------------------------------------------------------ |
| `message` | Required feedback text. Maximum length: 1024 characters.                                                     |
| `type`    | Optional feedback type. Use `ERROR` for corrections and `INFO` for notes. Defaults to `ERROR`.               |
| `fieldId` | Optional extracted field id. Include it when feedback applies to one field, or omit it for general feedback. |

## Related reference

* [Document Insights API](/api/api-reference/document-insights)
* [Event Reference](/api/event-reference)