---
title: "Document automation: from PDF to usable data | SUMāTO"
description: How document extraction is designed, why accuracy per field matters more than accuracy per document, and where the human review step belongs in it.
image: https://sumatogroup.com/hubfs/BRANDING/SUM%C4%81TO%20%7C%20LOGO%201000x500.png
---

[Skip to content](https://sumatogroup.com/en/insights/blog/automatizacion-documentos-pdf-dato#main-content)

- [INSIGHTS](https://sumatogroup.com/en/insights)
- [SUPPORT](https://sumatogroup.com/en/support)
- [CONTACT](https://sumatogroup.com/en/contact)

EN

[Español](https://sumatogroup.com/insights/blog/automatizacion-documentos-pdf-dato) [English](https://sumatogroup.com/en/insights/blog/automatizacion-documentos-pdf-dato)

[![SUMāTO Group — home](https://sumatogroup.com/hs-fs/hubfs/BRANDING/SMT%20-%20LOGO.png?width=40&height=40&name=SMT%20-%20LOGO.png)](https://sumatogroup.com/en)

- [HOME](https://sumatogroup.com/en/)
- About
  
  #### SUMāTO
  
    - [About us→](https://sumatogroup.com/en/about-us)
    - [Terms→](https://sumatogroup.com/en/legal)
    - [Legal→](https://sumatogroup.com/en/legal)
    - [Cookies→](https://sumatogroup.com/en/legal)
    - [Data protection→](https://sumatogroup.com/en/legal)
  
  
  #### METHODOLOGIES
  
    - [Design Thinking→](https://sumatogroup.com/en/methodologies#design-thinking)
    - [Lean Startup→](https://sumatogroup.com/en/methodologies#lean-startup)
    - [PMI→](https://sumatogroup.com/en/methodologies#pmi)
    - [Scrum→](https://sumatogroup.com/en/methodologies#scrum)
  
  
  #### Vendors
  
    - [AWS→](https://sumatogroup.com/en/vendors#aws)
    - [Cisco→](https://sumatogroup.com/en/vendors#cisco)
    - [Dahua→](https://sumatogroup.com/en/vendors#dahua)
    - [Fortinet→](https://sumatogroup.com/en/vendors#fortinet)
    - [Huawei→](https://sumatogroup.com/en/vendors#huawei)
    - [Microsoft→](https://sumatogroup.com/en/vendors#microsoft)
    - [OCI→](https://sumatogroup.com/en/vendors#oci)
    - [Panduit→](https://sumatogroup.com/en/vendors#panduit)
- Capabilities
  
  #### TECHNOLOGY
  
    - [Artificial Intelligence→](https://sumatogroup.com/en/artificial-intelligence)
    - [Data Analytics→](https://sumatogroup.com/en/data-analytics)
    - [Automation→](https://sumatogroup.com/en/automation-rpa)
    - [Cybersecurity→](https://sumatogroup.com/en/cybersecurity)
    - [Cloud→](https://sumatogroup.com/en/cloud)
  
  
  #### SEGMENTS
  
    - [SMB→](https://sumatogroup.com/en/smb)
    - [Enterprise→](https://sumatogroup.com/en/enterprise)
    - [Government→](https://sumatogroup.com/en/government)
- Consulting
  
  #### Assessments
  
    - [AI Readiness→](https://sumatogroup.com/en/ai-readiness-assessment)
    - [Analytics→](https://sumatogroup.com/en/data-analytics-maturity-assessment)
    - [Cloud→](https://sumatogroup.com/en/cloud-readiness-assessment)
    - [Cybersecurity→](https://sumatogroup.com/en/cybersecurity-assessment)
    - [Enterprise Architecture→](https://sumatogroup.com/en/enterprise-architecture-assessment)
    - [IT Maturity→](https://sumatogroup.com/en/it-maturity-assessment)
    - [IT Strategy→](https://sumatogroup.com/en/technology-strategy-assessment)
    - [Process Automation→](https://sumatogroup.com/en/process-automation-assessment)
  
  
  #### Consulting & Architecture
  
    - [AI First→](https://sumatogroup.com/en/ai-first)
    - [BCP→](https://sumatogroup.com/en/business-continuity-plan)
    - [DRP→](https://sumatogroup.com/en/disaster-recovery-plan)
    - [Enterprise Architecture→](https://sumatogroup.com/en/enterprise-architecture-togaf)
    - [Enterprise Transformation→](https://sumatogroup.com/en/enterprise-transformation)
    - [IT Strategic Plan→](https://sumatogroup.com/en/it-strategic-plan)
    - [Strategic Consulting→](https://sumatogroup.com/en/strategic-consulting)
- Operations
  
  #### INFRASTRUCTURE
  
    - [Data Center→](https://sumatogroup.com/en/data-center)
    - [Managed Services→](https://sumatogroup.com/en/managed-services)
    - [VDI→](https://sumatogroup.com/en/vdi)
    - [Intelligent Video Surveillance→](https://sumatogroup.com/en/video-surveillance)
  
  
  #### SECURITY
  
    - [NOC→](https://sumatogroup.com/en/noc)
    - [SOC→](https://sumatogroup.com/en/soc)
  
  
  #### USERS
  
    - [Modern Desktop→](https://sumatogroup.com/en/modern-desktop)
    - [Help Desk→](https://sumatogroup.com/en/help-desk)
- Industries
  
  Industries
  
    - [Banking & Finance→](https://sumatogroup.com/en/banking-finance)
    - [Insurance→](https://sumatogroup.com/en/insurance)
    - [Government→](https://sumatogroup.com/en/government)
    - [Healthcare→](https://sumatogroup.com/en/healthcare)
    - [Telecommunications→](https://sumatogroup.com/en/telecommunications)
    - [Retail & Consumer→](https://sumatogroup.com/en/retail)
    - [Manufacturing→](https://sumatogroup.com/en/manufacturing)
    - [Energy, Oil & Gas→](https://sumatogroup.com/en/energy-oil-gas)
    - [Education→](https://sumatogroup.com/en/education)
    - [Logistics & Transportation→](https://sumatogroup.com/en/logistics-transport)
    - [Legal Services→](https://sumatogroup.com/en/legal-services)
    - [Engineering & Construction→](https://sumatogroup.com/en/engineering-construction)
- Resources
  
  #### CONTENT
  
    - [Blog→](https://sumatogroup.com/en/insights)
    - [Use cases→](https://sumatogroup.com/en/use-cases)
  
  
  #### EVENTS
  
    - [Webinars→](https://sumatogroup.com/en/webinars)

EN

[Español](https://sumatogroup.com/insights/blog/automatizacion-documentos-pdf-dato) [English](https://sumatogroup.com/en/insights/blog/automatizacion-documentos-pdf-dato)

Search

- There are no suggestions because the search field is empty.

[automatizacion](https://sumatogroup.com/en/insights/tag/automatizacion)

# Document automation: turning a PDF into data you can rely on

[Andrés Lozada](https://sumatogroup.com/en/insights/author/andres-lozada) · Apr 18, 2023, 7:00:00 AM · 8 min read

**Extracting data from a document is not one problem. It is three — finding the document, reading the fields, and deciding whether to trust what was read — and most projects that disappoint spent their effort on the second while the failures came from the third.**

Below: which documents are worth automating, why per-field accuracy is the number that matters, where human review belongs, and what to measure before choosing anything.

Invoices, delivery notes, bank statements, identity documents, contracts. Every organisation has a queue of these arriving as PDFs, images or scans, and someone re-typing them into a system.

The technology to read them is no longer the constraint. The design around it is.

That shift is recent enough to have changed which questions matter. Five years ago the sensible first question was whether a document could be read at all; today it is what happens to the documents that are read wrongly, and how you find out.

## Which documents are worth automating

The variables that decide this are volume, structure and consequence. High volume with stable structure is the obvious case. Low volume with high consequence — a contract clause, a compliance filing — is usually not, because the review a human has to do anyway is most of the work.

Structure matters more than format. A PDF issued by one supplier in the same layout every month is easy regardless of whether it is text or an image. A hundred suppliers each with their own layout is a different problem, and the one where generic tools disappoint.

The useful screening question is how many distinct layouts the queue actually contains. Teams usually guess low. Counting them on a real month of documents is an afternoon and reshapes the estimate more than any other input.

## Per field, not per document

A vendor accuracy figure of ninety-five per cent means very little without knowing which field. A document with twelve fields at ninety-five per cent each has a low chance of being entirely correct, and the fields do not matter equally.

What matters is accuracy on the fields that carry consequence — the amount, the account, the identifier — and the error mode on each. A field that fails by leaving itself empty is safe: it stops. A field that fails by returning a confident wrong value is the one that reaches the ledger.

So the specification is per field, with a required confidence and a defined behaviour below it. That specification is what a supplier should be evaluated against, and it is almost never how the evaluation is run.

## Where the human review step belongs

Not at the end, reviewing everything. That reproduces the manual process with an extra step. Review belongs exactly where confidence is below the threshold for a consequential field, and nowhere else.

Designed that way, the reviewer sees a small fraction of documents, sees only the fields in doubt, and sees the source image beside the extracted value. That last detail decides whether review takes eight seconds or two minutes, and at volume it is the difference between viable and not.

The review decisions should also feed back. A field corrected the same way repeatedly is a pattern the extraction should learn, and an estate that never closes that loop keeps the same error rate indefinitely.

## The part that is not extraction

Finding the document is frequently harder than reading it. Documents arrive by email, through a portal, on paper, attached to a message thread with three other files, or embedded inside a forwarded chain.

Classifying what arrived — which of these is the invoice, which is the delivery note, which is an unrelated attachment — is a step in its own right, and projects that skipped it discover it in testing.

It is worth scoping deliberately, because the answer differs by source. A supplier portal delivering one file type on a schedule needs almost nothing; a shared mailbox receiving everything from everyone needs a classification step that is a small project of its own.

The same is true at the other end. Extracted data has to be written somewhere, matched against an existing record, and reconciled when it does not match. That downstream half is usually where the remaining manual effort survives.

## Identity documents deserve their own answer

Reading a national identity document is technically the easiest case in this article — fixed layout, small field count, mature tooling — and operationally the most constrained, because the data extracted is personal data from the moment it is read.

That shapes the design rather than adding a compliance review at the end. Where the images are stored and for how long, who can see the extracted fields, whether the document image is kept at all once the fields are validated, and what a supplier processing it on your behalf is permitted to retain — these are design decisions, and they are read against Habeas Data in Colombia and the federal data protection law in Mexico.

The practical rule that avoids most of the difficulty: extract, validate, discard the image, and keep only the fields the process actually needs. Retention that has no purpose is the part that turns a routine capability into an exposure.

## What to measure first

Three numbers, all available before choosing any tool: documents per month, distinct layouts in a real sample, and the current time to process one by hand including the exceptions.

A fourth is worth adding where it applies: how often a document currently gets processed wrong and has to be corrected. That figure sets the bar the automation has to beat, and it is often lower than the team assumes, which changes what accuracy is acceptable.

These four size the case honestly. In Colombia and Mexico they also determine whether it is a case at all, because the manual cost being displaced is lower than the vendor pricing usually assumes.

## The failure that gets discovered late

Documents change. A supplier redesigns its invoice, a bank adds a column to its statement, a form gets a new field. Extraction that was tuned to a layout degrades quietly rather than breaking loudly.

The defence is monitoring the extraction rate per source over time, not just the error queue. A supplier whose confidence scores dropped last month is a signal available weeks before anyone notices the data is wrong.

The same monitoring answers a question that comes up in every renewal conversation: whether the extraction is getting better or worse. Without a trend per source, that discussion is conducted on impressions, and impressions are formed by whoever handled the last bad document.

Without that, the discovery route is a finance query about a figure that looks off, three weeks after the layout changed, with a month of records to re-check.

## Where to start

With one document type, one source, and the four numbers above. A first implementation scoped to the highest-volume layout proves the design and produces the measurements the rest of the case needs.

A [process automation assessment](https://sumatogroup.com/en/process-automation-assessment) establishes those and calculates the return with local costs, so the [automation](https://sumatogroup.com/en/automation-rpa) plan starts from the document queue that actually justifies it.

## Frequently asked questions

### Which documents are worth automating?

High volume with stable structure. Low volume with high consequence usually is not, because the human review needed anyway is most of the work. Count the distinct layouts before estimating.

### Why is per-field accuracy the number that matters?

Because a document with twelve fields at ninety-five per cent each is rarely entirely correct, and the fields do not matter equally. The specification should be per field with a required confidence.

### Which extraction errors are dangerous?

A field that fails by returning a confident wrong value, not one that fails by leaving itself empty. An empty field stops; a wrong one reaches the ledger.

### Where should human review sit?

Only where confidence falls below the threshold on a consequential field, showing the reviewer the source image beside the extracted value. Reviewing everything reproduces the manual process.

### What is harder than the extraction itself?

Finding and classifying the document on arrival, and matching the extracted data to an existing record downstream. Both are where the surviving manual effort tends to be.

### How do you catch a layout change?

By monitoring extraction confidence per source over time, not only the error queue. A drop shows up weeks before anyone notices the resulting data is wrong.

Next step

[Automation and RPA](https://sumatogroup.com/en/automation-rpa)[Process Automation Assessment](https://sumatogroup.com/en/process-automation-assessment)[Data Analytics](https://sumatogroup.com/en/data-analytics)

[automatizacion](https://sumatogroup.com/en/insights/tag/automatizacion)

![Andrés Lozada](https://sumatogroup.com/hs-fs/hubfs/SPEAKERS/AL.jpeg?width=56&height=56&name=AL.jpeg)

Andrés Lozada Apr 18, 2023, 7:00:00 AM 

[LinkedIn](https://www.linkedin.com/in/andreslozada/)

### Explore more from SUMāTO

[Enterprise AI](https://sumatogroup.com/en/artificial-intelligence) [Enterprise Transformation](https://sumatogroup.com/en/enterprise-transformation) [Strategic Consulting](https://sumatogroup.com/en/strategic-consulting) [AI Agent](https://sumatogroup.com/en/artificial-intelligence) [AI Contact Center](https://sumatogroup.com/en/artificial-intelligence) [Cybersecurity](https://sumatogroup.com/en/cybersecurity)

### Related Posts

#### [Intelligent Document Management with AI: From Document Chaos to Structured Knowledge](https://sumatogroup.com/en/insights/blog/gestion-documental-inteligente-conocimiento-estructurado)

Nearly every modern organization faces a paradox: never has so much information been available, and never has it been so hard to find exactly what...

#### [From Chatbots to Autonomous Agents: The Evolution Your Company Can't Ignore](https://sumatogroup.com/en/insights/blog/chatbots-agentes-autonomos-evolucion-ia-empresarial)

Between 2018 and 2022, virtually every company with more than 500 employees in Latin America deployed some variant of a chatbot. The projects were...

#### [AI Agents: From Copilot to Autonomous Agent](https://sumatogroup.com/en/insights/blog/agentes-de-ia)

Over the past year, copilots have become familiar: a window that suggests code, drafts an email, or summarizes a document while you keep control of...

![SUMāTO](https://sumatogroup.com/hs-fs/hubfs/BRANDING/SUM%C4%81TO%20%7C%20LOGO%201000x500.png?width=200&height=100&name=SUM%C4%81TO%20%7C%20LOGO%201000x500.png)

Strategic technology planning consultants.

AI, Analytics, Cloud and Cybersecurity

<https://www.linkedin.com/company/sumatogroup> <https://www.youtube.com/@sumatogroup>

## Navigation

[Home](https://sumatogroup.com/en) [Capabilities](https://sumatogroup.com/en/artificial-intelligence) [Consulting](https://sumatogroup.com/en/strategic-consulting) [Operations](https://sumatogroup.com/en/managed-services) [Industries](https://sumatogroup.com/en/banking-finance) [Resources](https://sumatogroup.com/en/insights)

## SUMāTO

[About](https://sumatogroup.com/en/about-us) [Terms](https://sumatogroup.com/en/legal#terminos) [Legal & Privacy](https://sumatogroup.com/en/legal) [Data protection](https://sumatogroup.com/en/legal)

Cookies

## [Contact](https://sumatogroup.com/en/contact)

[sales@sumatogroup.com](mailto:sales@sumatogroup.com)

Mexico HQ

Mexico City, Mexico

[+52 55 8897 5791](tel:+525588975791)

Bogotá

Bogotá, Colombia

[+57 601 724 5059](tel:+576017245059)

© 2026 SUMāTO Group. All rights reserved.

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Andrés Lozada",
    "url" : "https://sumatogroup.com/en/insights/author/andres-lozada"
  },
  "datePublished" : "2023-04-18T13:00:00.000Z",
  "headline" : "Document automation: from PDF to usable data | SUMāTO",
  "mainEntityOfPage" : {
    "@id" : "https://sumatogroup.com/en/insights/blog/automatizacion-documentos-pdf-dato",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://sumatogroup.com/hubfs/BRANDING/Logo_SUMATO_Original%20-%201000x500.png"
    },
    "name" : "SUMāTO Group"
  }
}
```

```json
{
  "@context" : "https://schema.org",
  "@type" : "FAQPage",
  "mainEntity" : [ {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "High volume with stable structure. Low volume with high consequence usually is not, because the human review needed anyway is most of the work. Count the distinct layouts before estimating."
    },
    "name" : "Which documents are worth automating?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "Because a document with twelve fields at ninety-five per cent each is rarely entirely correct, and the fields do not matter equally. The specification should be per field with a required confidence."
    },
    "name" : "Why is per-field accuracy the number that matters?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "A field that fails by returning a confident wrong value, not one that fails by leaving itself empty. An empty field stops; a wrong one reaches the ledger."
    },
    "name" : "Which extraction errors are dangerous?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "Only where confidence falls below the threshold on a consequential field, showing the reviewer the source image beside the extracted value. Reviewing everything reproduces the manual process."
    },
    "name" : "Where should human review sit?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "Finding and classifying the document on arrival, and matching the extracted data to an existing record downstream. Both are where the surviving manual effort tends to be."
    },
    "name" : "What is harder than the extraction itself?"
  }, {
    "@type" : "Question",
    "acceptedAnswer" : {
      "@type" : "Answer",
      "text" : "By monitoring extraction confidence per source over time, not only the error queue. A drop shows up weeks before anyone notices the resulting data is wrong."
    },
    "name" : "How do you catch a layout change?"
  } ]
}
```

```json
{
  "@context" : "https://schema.org",
  "@id" : "https://sumatogroup.com/#organization",
  "@type" : "Organization",
  "address" : {
    "@type" : "PostalAddress",
    "addressCountry" : "MX",
    "addressLocality" : "Huixquilucan",
    "addressRegion" : "Estado de México",
    "postalCode" : "52787",
    "streetAddress" : "Av. Vialidad de la Barranca No. 6, Torre 1, Suite 400, Piso 4, Col. Bosques de las Palmas"
  },
  "alternateName" : [ "SUMāTO Group", "SUMATO Group", "Sumato Group", "SUMATO", "SUMaTO", "SUMaTO Group", "SUMTO", "SUMTO Group" ],
  "areaServed" : [ {
    "@type" : "Country",
    "name" : "México"
  }, {
    "@type" : "Country",
    "name" : "Colombia"
  }, {
    "@type" : "Place",
    "name" : "Latinoamérica"
  } ],
  "contactPoint" : {
    "@type" : "ContactPoint",
    "areaServed" : "Latinoamérica",
    "availableLanguage" : [ "es", "en" ],
    "contactType" : "sales",
    "email" : "sales@sumatogroup.com"
  },
  "description" : "SUMāTO is a Latin American technology consulting and integration firm founded in 2016, with a presence in Mexico and Colombia. It designs, implements and operates artificial intelligence, data analytics, automation, cybersecurity and cloud on the systems a client already runs, under governance frameworks such as NIST AI RMF and ISO/IEC 42001.",
  "foundingDate" : "2016",
  "knowsAbout" : [ "Inteligencia Artificial", "IA Generativa", "Agentes de IA", "Analítica de Datos", "Big Data", "Automatización de Procesos (RPA)", "Ciberseguridad", "Computación en la Nube", "Continuidad del Negocio y Recuperación ante Desastres", "Arquitectura Empresarial", "Transformación Digital" ],
  "legalName" : "SUMāTO Group",
  "location" : [ {
    "@type" : "Place",
    "address" : {
      "@type" : "PostalAddress",
      "addressCountry" : "MX",
      "addressLocality" : "Huixquilucan",
      "addressRegion" : "Estado de México",
      "postalCode" : "52787",
      "streetAddress" : "Av. Vialidad de la Barranca No. 6, Torre 1, Suite 400, Piso 4, Col. Bosques de las Palmas"
    },
    "name" : "SUMāTO MX",
    "telephone" : "+52 55 8897 5791"
  }, {
    "@type" : "Place",
    "address" : {
      "@type" : "PostalAddress",
      "addressCountry" : "CO",
      "addressLocality" : "Bogotá",
      "streetAddress" : "Cra. 45 # 103-34, Of. 202"
    },
    "name" : "SUMāTO CO",
    "telephone" : "+57 601 724 5059"
  } ],
  "logo" : {
    "@type" : "ImageObject",
    "height" : 500,
    "url" : "https://sumatogroup.com/hubfs/BRANDING/SUM%C4%81TO%20%7C%20LOGO%201000x500.png",
    "width" : 1000
  },
  "name" : "SUMāTO",
  "sameAs" : [ "https://www.linkedin.com/company/sumatogroup", "https://www.youtube.com/@sumatogroup", "https://torre.ai/teams/SUMaTOGroup", "https://www.cbinsights.com/company/sumto-group", "https://elioplus.com/profiles/channel-partners/57295/sumato-group" ],
  "telephone" : "+52 55 8897 5791",
  "url" : "https://sumatogroup.com"
}
```

```json
{
  "@context" : "https://schema.org",
  "@id" : "https://sumatogroup.com/#website",
  "@type" : "WebSite",
  "description" : "Technology consulting in AI, data, automation, cybersecurity and cloud across Latin America.",
  "inLanguage" : "en",
  "name" : "SUMāTO",
  "publisher" : {
    "@id" : "https://sumatogroup.com/#organization"
  },
  "url" : "https://sumatogroup.com"
}
```

```json
{
  "@context" : "https://schema.org",
  "@type" : "BreadcrumbList",
  "itemListElement" : [ {
    "@type" : "ListItem",
    "item" : "https://sumatogroup.com/en",
    "name" : "Home",
    "position" : 1
  }, {
    "@type" : "ListItem",
    "item" : "https://sumatogroup.com/en/insights/blog/automatizacion-documentos-pdf-dato",
    "name" : "Document automation: turning a PDF into data you can rely on",
    "position" : 2
  } ]
}
```