Transforming Extraction Results

Run your own rules on extraction results - reformatting, unit conversion, lookups and calculated fields

What a transformation is

A transformation is a small rule that runs on a Standardization after DocuPipe extracts it. Use it for logic that should give the same answer every time: reformatting names, converting units, adding a field calculated from other fields, filling in a value from a lookup table, or reshaping the output for the system that receives it.

The original extraction is never changed. The transformed version is stored next to it, so you can always compare the two.

📘

A transformation's output does not have to match your Schema. It can add fields, rename them, or return a completely different shape.

Creating a transformation

  1. Open the Transformation page from the left menu and click + Transformation.
  2. Describe the rule in plain words. For example: "Names are written as Last, First. Rewrite every name as First Last."
  3. Optionally give it a name, and pick a sample to test it on: a standardization, or a CSV, Excel, JSON or XML document.
  4. Click Write Transformation.

We write the code, run it on your sample, and review it against your description. It is saved only if it passes the review. The result shows the code, the review verdict with its reasons, and the output on your sample.

You can close the window while it works. Clicking + Transformation again brings you back to the progress or the result.

👍

Always pick a sample when you can. You see the real output before relying on the rule, and the review checks the code against what it actually did.

Changing a transformation

Select a transformation in the table and open Actions, or open it and use More Actions:

  • Describe a change - explain what should be different, and we rewrite the code.
  • Edit code - edit the Python yourself. Your code is checked, run on your sample, and reviewed the same way.

A change is saved as a new transformation. The original stays as it is, and the new one shows which transformation it was edited from.

Running transformations

From the dashboard: in the Standardize window, pick one or more under Transformations (Optional) in the configure step. They run in the order you choose, each one receiving the previous one's output. A schema must be selected.

From the API: pass transformationIds to POST /v3/standardize or POST /v2/standardize/batch, in the order they should run. The standardization completes and its webhook fires first; the transformations then run as their own job, and the standardization job carries its transformationJobId. You receive transformation.processed.success or transformation.processed.error when it finishes. See Receiving Results with Webhooks.

Where the result appears

In the dashboard, a Transformed tab appears in the standardization window once a transformed result exists, with JSON and Excel downloads. The Excel file follows the field order of the transformed result. The Viewer tab still shows the original extraction.

In the API, the standardization's transformedData holds the result.

Using the API

Everything on the Transformation page is also available through the API.

Create or change a transformation: POST /transformation with a description, and optionally a name and one sample (standardizationId, or documentId of a CSV, Excel, JSON or XML document). To use your own code, add code. To change an existing transformation, add basedOnTransformationId with either a description of the change or the edited code. The call returns a jobId; poll GET /job/{job_id} until status is no longer processing. The completed job has outcome (approved or rejected), the code, an explanation, the review's verdicts and reasons, the output on your sample in preview, and the credits charged. When approved, transformationId is set.

List or fetch transformations: GET /transformations lists them, newest first, with their code. GET /transformation/{transformation_id} returns one.

Run transformations on an existing result: POST /transformation/apply with transformationIds and exactly one of standardizationId or documentId. The transformed result is returned and stored on that standardization or document. To transform as part of a new extraction instead, pass transformationIds to POST /v3/standardize as described above.

Fetch a stored result: GET /transformation/result/{standardization_id} or GET /transformation/document-result/{document_id}.

📘

A transformation's code is checked before anything runs. If your own code uses something that is not allowed, POST /transformation returns a 400 explaining why, and nothing is charged.

Writing your own code

A transformation is a Python function:

def transform(data, context):
    data["fieldCount"] = len(data)
    return data
  • data is the extraction result (or, for a spreadsheet or data file used as a sample, its contents).
  • Return anything that can be written as JSON.
  • Only standard-library modules for working with data can be imported, such as re, math, datetime, decimal, json, collections, itertools, statistics, csv and string. Code that reads files, reaches the network or evaluates other code is rejected.
  • The code runs in an isolated environment with no network access and only sees the data it was given.
  • Each run has a time limit of 60 seconds, and up to 50 transformations can be chained in one run.
🚧

If your code raises an error while running on a document, the original extraction is kept, the transformed result is not updated, and the run is still charged.

Cost

OperationCredits
Creating or changing a transformation3 per request, charged once it is approved or rejected
Running transformations1 per document, however many are chained

Submitting your own code that fails the code checks or the test run on your sample is free. See Understanding Credits and Billing.


Did this page help you?