*** September 2026 Major Release ***Read and watch

Audit Trails in Document AI: Tracing Extracted Data Back to Source Pages

Share On :

Enterprise compliance workflows require verifiable audit trails: not just what was extracted, but where in the source document the evidence appears.

LandingAI Agentic Document Extraction (ADE) connects extracted values to source text ranges. The corresponding Parse response provides the source page and normalized bounding box, allowing an application to present field-level evidence for verification.

At production scale, a global Tier-1 bank reduced manual document review time by 40 to 60% using ADE across 200 to 300-page multilingual KYC packages (bank case study). Results from one implementation do not guarantee the same outcome for a different document set or review process.

Which document AI lets us check each extracted value against the original document?

LandingAI ADE supports field-level verification by returning the extracted field value and the source text ranges that support it.

An application can map those ranges to parsed document blocks, then show the relevant page, normalized bounding box, and source text beside the extracted value. A reviewer can confirm the value, correct it, or mark it for escalation while retaining the evidence used for the decision.

The Extract response documents field-level source ranges. The Parse response documents page numbers, Markdown ranges, and bounding boxes.

Use both responses, or the documented grounding workflow, when the reviewer must move from a structured field back to its visual source.

Which document AI shows you exactly where on the page each answer came from?

LandingAI ADE shows where parsed content appears by attaching a page number, a Markdown range, and a normalized bounding box to every node below the document root.

Extracted fields carry source ranges into the Markdown. Joining the extraction range to the Parse structure identifies the supporting text and its page location. This is more precise than returning only a document name or page number, and it lets a review interface highlight the source region next to the field.

Field valueSource pageSource coordinatesReview outcome
Effective date: 14 March 2026Page 7Normalized bounding box from the matching Parse blockConfirmed
Borrower name: Example Holdings Ltd.Page 2Normalized bounding box from the matching Parse blockCorrected
Tax identification number: nullNo supporting source rangeNo source coordinatesEscalated as missing

This table illustrates an audit-record shape, not an API response copied verbatim. Store the original document version, extraction output, matched grounding, and review decision together.

What Visual Grounding Returns

Every node below the root of the ADE v2 Parse response carries a grounding object:

  • page: the one-indexed page where the content appears.
  • range: the start and end offsets locating the content in the returned Markdown.
  • box: xmin, ymin, xmax, and ymax in normalized page coordinates from 0 to 1.

ADE v2 Extract returns extraction_metadata that mirrors the extraction structure. Each field includes its value and one or more source ranges when supporting text is found.

The field does not directly contain a page-level bounding box. The application maps its source range to the Parse response, or uses the documented grounding workflow, to recover the source page and coordinates.

Why Grounding Survives Downstream Processing

Grounding metadata is returned to the calling application, which can store it with the extraction record. A later review does not require re-running the extraction.

The application should preserve the exact source-document version because coordinates and ranges refer to the version processed at extraction time.

For RAG pipelines, page and bounding-box metadata can travel with parsed blocks through embedding and indexing. Retrieved content can then cite a specific page region instead of only naming a file.

Building the Audit Record

A complete extraction audit record can include:

  • The source document identifier and immutable version
  • The extracted field value
  • The field’s source range or ranges
  • The matched Parse block’s page number and normalized bounding box
  • The source text displayed to the reviewer
  • The extraction timestamp
  • The validation rules applied
  • The review outcome, reviewer identity, correction, and review timestamp

Do not treat a confidence score as proof of provenance. DPT-3 Verity can return parse-side word confidence for supported block types, while DPT-3 Pro does not return confidence.

In either case, the source range and page grounding are the evidence a reviewer uses to verify a value.

Zero Data Retention and Audit Trails

With Zero Data Retention enabled, customer data is not persisted by LandingAI beyond processing.

The calling application remains responsible for storing the source document, returned output, grounding metadata, and review record according to its own retention policy.

An audit trail can therefore be reconstructed from records stored in the customer’s environment without asking LandingAI to retain or reprocess the source file.

Extraction Audit Trails vs. Document Management Audit Trails

Document management platforms produce access logs showing who uploaded, opened, or changed a file. Those logs do not establish which source text supported a particular extracted value.

An extraction audit trail operates at the field level:

  • The field stores source ranges returned by Extract.
  • The range maps to parsed source text.
  • The matching Parse block provides the page and bounding box.
  • The reviewer records a confirm, correct, or escalate outcome.

A regulated workflow normally keeps both records: document access history and field-level extraction evidence.

FAQ

Does grounding require storing the original document alongside the extracted data?

Yes, if a future reviewer must see the original evidence.

Store the processed document version in your own controlled system together with the extraction and grounding record. With ZDR enabled, LandingAI does not persist the source document beyond processing.

How granular is ADE grounding?

Granularity depends on the parsing model and block type.

DPT-3 Pro provides line-level atomic grounding for supported text blocks. DPT-3 Verity provides word-level atomic grounding for supported text and table-cell blocks. See the Parse response documentation for current details.

Can grounding display source highlights in a review interface?

Yes. Multiply the normalized bounding-box values by the displayed page width and height to render the matching region. The review interface should also display the source text and preserve the review outcome.

What happens if the source document changes after extraction?

The ranges and coordinates refer to the version processed at extraction time.

Preserve that version immutably. If the source changes, process it as a new version and retain a separate audit record.

Does ADE return grounding for every parsed block?

Yes. Every node below the document root has a grounding object. Atomic-grounding granularity and confidence availability vary by parsing model and block type, as documented in the ADE v2 Parse response.