← All articles
EngineeringDataWorkflows

From Unstructured Data to Structured Data: Building AI Data Pipelines for Banking and Healthcare

Turn faxes, PDFs, documents, images, forms and free-form text into structured, validated data, and send it straight into the systems that need it. Extraction is one node in a workflow, not the whole job.

Most production systems do not operate on documents. They operate on structure.

A loan processing system expects fields like applicant_name, annual_income, account_balance and loan_amount. A healthcare workflow might expect patient_name, date_of_birth, provider, diagnosis, insurance_information and referral_reason.

But the information entering these organisations rarely arrives that cleanly. It arrives inside PDFs, scanned documents, images, forms, clinical notes, bank statements, pay stubs, referral letters, insurance documents and free-form text. The information is there. The structure is not.

For years, organisations have closed that gap with some combination of OCR, document templates, regex, custom parsers, manual data entry and increasingly complex document-processing pipelines. Multimodal AI changes what is possible. But getting a model to understand a document is only the beginning.

How do you reliably turn unstructured information into exactly the structure your downstream system expects, and then do something with it?

The problem is not reading the document

Consider a bank statement. A multimodal model can look at it and explain what it contains: a statement for September, with account information, opening and closing balances, deposits, withdrawals and transaction history.

That is impressive. It is also not particularly useful to a production system. A lending workflow does not need a paragraph explaining the document. It needs fields.

What the lending workflow actually needs · JSON
{
  "account_holder": "John Smith",
  "account_type": "Checking",
  "statement_start_date": "2026-09-01",
  "statement_end_date": "2026-09-30",
  "opening_balance": 38240.12,
  "closing_balance": 42850.72,
  "total_deposits": 12450.00,
  "total_withdrawals": 7839.40,
  "currency": "USD"
}

That is a fundamentally different problem. We are moving from document understanding to schema-constrained extraction to doing something with the result.

And that last step matters. The structured result usually is not the final product. It needs to be validated, transformed, stored, passed to another API, reviewed by a person, or used to trigger another workflow.

Unstructured data
Extract
Validate
Transform
Act on it
Store
The shape of the problem once it has to run in production.

Why banking and healthcare make this obvious

Almost every industry deals with unstructured information, but banking and healthcare make the challenge particularly visible. Both operate sophisticated structured systems while continuously receiving information in highly variable formats.

Banking

A lending or financial workflow may need to process bank statements, loan applications, mortgage documents, pay stubs, tax documents, KYC documents, invoices, financial statements, transaction reports and scanned forms. The data eventually needs to become fields another system can understand.

Healthcare

Healthcare organisations encounter referral documents, patient intake forms, lab reports, clinical notes, insurance documents, discharge summaries, medical records, prescriptions, scanned forms and PDFs received from external providers.

And a great deal of it still arrives by fax. Fax remains one of the most common ways referrals and records move between providers in the United States, which means a large share of inbound clinical information lands as a scanned image of a printed page: no text layer, variable quality, no structure at all. A fax inbox is an unstructured input like any other, and it is often the highest-volume one.

Again, the destination is usually structured. The source is not. That mismatch creates an enormous amount of glue code and operational work.

From a document to a schema

Take a simplified banking example. A lending application receives three files: a bank statement, a pay stub and a loan application. Different documents, different layouts, potentially different institutions.

But the lending system does not care about their layouts. It cares about a defined set of information.

The contract the application defines · JSON
{
  "applicant": {
    "full_name": "",
    "address": ""
  },
  "employment": {
    "employer": "",
    "monthly_income": 0
  },
  "financials": {
    "account_balance": 0,
    "monthly_deposits": 0
  },
  "loan": {
    "requested_amount": 0,
    "purpose": ""
  }
}

The input can remain messy. The output cannot. Instead of asking “what does this document say?” we are asking “given this input, populate this exact structure.”

That turns multimodal intelligence into something conventional software can consume.

Building this in Xpectrum

In Xpectrum this is a workflow. Extraction is not an isolated AI call; it is one node in something that runs.

An Xpectrum workflow: trigger, parameter extractor, custom function, then an HTTP API request, a Mongo insert and a Supabase row, ending at the output
Trigger, parameter extractor, custom function, then whichever systems need the result. The trigger takes text, a document upload and an image upload as separate input fields. Open full size \u2197

The trigger panel is worth a look. The inputs are declared as fields the workflow accepts, so a paragraph of free text, an uploaded document and an uploaded image are all first-class entry points rather than special cases.

Step 1: accept unstructured input

The workflow starts with the data the organisation already has: text, an image, a document, a PDF, or another supported input.

A banking workflow might receive a loan application, a bank statement and a pay stub. A healthcare workflow might receive a referral PDF, an insurance document and a clinical note.

The workflow does not need every source to share the same structure before processing begins. That is the point.

Step 2: define what you actually want

This is the most important step. Instead of letting the model decide what information is interesting, the developer defines what the application needs.

A healthcare referral schema · JSON
{
  "patient": {
    "name": "string",
    "date_of_birth": "date"
  },
  "referral": {
    "referring_provider": "string",
    "specialty": "string",
    "reason": "string",
    "priority": "string"
  },
  "insurance": {
    "provider": "string",
    "member_id": "string"
  }
}
The parameter extractor configured with required fields: name, date of birth, insurance provider and member id, each with its own instruction
Each field is declared with its own instruction. Note the date of birth: “retrieve dob from the data in mm-dd-yyyy format always”, so the format is part of the contract rather than something to clean up later. Open full size \u2197

The extractor takes an input variable from the trigger and a list of parameters, each one named, typed and required or not. The instruction attached to a field is where you pin down the detail that would otherwise drift between documents.

The parameter extractor maps the unstructured input into those parameters. Your database does not need to understand a PDF. Your API does not need to interpret a clinical note. Your application receives the fields it was designed to consume.

Step 3: normalise the result

Extraction is rarely the end of the pipeline. Real documents are messy. A date may arrive as 09/14/26, September 14, 2026 or 14-Sep-2026. Currency may arrive as $12,450, USD 12,450.00 or 12.45K USD. Names, addresses, identifiers and categorical values all vary.

So after extraction, the result passes through a custom function. “September 14, 2026” becomes “2026-09-14”. “$12,450.00” becomes 12450.00. One source’s “cardiac specialist” becomes the “Cardiology” the receiving system expects.

AI handles ambiguity. Code handles deterministic business logic.

That is a better production pattern than asking the model to own every step.

Step 4: validate before anything downstream

Extraction should not automatically imply trust. Suppose a mortgage workflow extracts a monthly income, an account balance and a requested loan amount. Before that reaches another system, the workflow may need to check whether required fields are present, whether numeric values are actually numeric, whether dates are valid, whether expected identifiers exist, whether the result satisfies application-specific rules, and whether the case should go to a person.

This is where production AI starts to look different from an AI demo. The objective is not to make the model answer. It is to make the system behave predictably.

LangChain makes a similar distinction in its production-agent writing: model intelligence is one component, and production systems need the surrounding runtime, controls, observability and operational plumbing. LangChain ↗ The same principle applies here. Extraction is one capability inside a larger system.

Step 5: put the data where work happens

Once information has been extracted and transformed, it should not stop inside the AI platform. It becomes input to the next system.

Structured data
MongoDB
Application
Supabase
Database
HTTP API
Existing system
AI stops being an assistant beside the workflow and becomes part of the data pipeline.

Example: bank statement processing

A customer submits a bank statement. The raw input might contain customer information, account numbers, the statement period, balances, deposits, withdrawals, transactions, fees and a good deal of boilerplate.

The application may only need a fraction of it.

Extracted financial fields · JSON
{
  "account_holder": "John Smith",
  "statement_period": {
    "from": "2026-09-01",
    "to": "2026-09-30"
  },
  "opening_balance": 38240.12,
  "closing_balance": 42850.72,
  "total_deposits": 12450.00,
  "total_withdrawals": 7839.40
}
Bank statement
Parameter extraction
Structured financial fields
Validation and normalisation
Internal API or database
Statements from different institutions differ. The output contract does not have to.

Example: healthcare referral intake

A referral might arrive as a fax or a PDF containing patient demographics, insurance information, provider information, medical context and the reason for referral. The receiving system may only need a handful of fields.

Extracted referral record · JSON
{
  "patient_name": "Jane Doe",
  "date_of_birth": "1982-04-16",
  "referring_provider": "Dr. Smith",
  "specialty": "Cardiology",
  "reason_for_referral": "Cardiac evaluation",
  "insurance_provider": "Example Health",
  "priority": "Routine"
}
Referral document
Extract required parameters
Normalise values
Validate required fields
Human review, if required
Create structured record
Send to downstream system
The AI is not replacing the workflow. It is removing the boundary between the document and the software that needs what is inside it.

Note what the extracted record contains: a name, a date of birth, a member id. That is protected health information, which makes the next section less optional than it looks.

Security is part of the pipeline, not a wrapper around it

Document pipelines tend to carry the most sensitive material an organisation holds. A referral contains a name, a date of birth and a member id. A bank statement contains an account number and a full transaction history. A KYC packet contains identity documents.

Extraction does not reduce that exposure. It concentrates it. One workflow now reads every document, holds the fields it pulled out, and writes them into several systems. That is a single place worth getting right, and a single place worth auditing.

What the architecture has to account for

  • Scoped credentials. The workflow should act with permissions that belong to the caller or to the workflow itself, not with a blanket key that can reach every record in the system.
  • A narrow blast radius for the model. The extractor needs the document and the schema. It does not need standing access to the database it eventually writes to.
  • An audit trail. Which document was processed, which fields came out, which system received them, and who approved the ones that needed approval.
  • Retention limits. The source document, the intermediate extraction and the logs all persist somewhere. Each needs a deliberate lifetime rather than an accidental one.
  • Encryption in transit and at rest, including the intermediate state between extraction and the write, which is easy to overlook because it feels transient.
  • Human review as a control, not just a quality step. For some categories the right answer is that nothing reaches a system of record without a person signing it off.

For workflows touching electronic protected health information in the United States, the HIPAA Security Rule requires administrative, physical and technical safeguards for ePHI. Fax makes this concrete: the moment a referral lands in an inbound fax queue it is already PHI, before any model has looked at it, which means the queue, the storage behind it and the processing that follows all sit inside the same obligation.

Document intelligence and compliance architecture are separate requirements. Solving the first does not satisfy the second.

The useful consequence of putting extraction inside a workflow is that these controls live at the workflow layer, where they apply to every document the same way, rather than being reimplemented per integration.

Extraction is only one node

This is perhaps the most important design decision. Unstructured-to-structured extraction should not be a standalone application. It should be a primitive inside a larger workflow.

PDF to extract to validate to API is useful. But real business processes get more interesting quickly.

Document uploaded
Extract parameters
Validate
Complete: continue
Incomplete: human review
Call internal API
Create database record
Trigger next workflow
At this point you are not building a document parser. You are building an operational system.

One intelligence layer, deterministic logic around it

Models are valuable precisely because the inputs are not predictable. Traditional software works extremely well when field A is here, field B is there, and field C always follows a known format. Real documents do not behave that way.

AI gives us a flexible interpretation layer. But once the document has been interpreted, much of the remaining process can, and often should, be deterministic.

Unstructured world
AI interpretation
Structured boundary
  • Validation
  • Business rules
  • Transformation
  • Permissions
Systems of record
AI handles what traditional software struggles with. Traditional software handles what it does best.

What happens when extraction is not certain?

Production extraction should not pretend uncertainty does not exist. Documents can be incomplete. Scans can be poor. Two values can conflict. A required field may not exist. A document might not even be the type you expected.

For high-impact workflows, the right architecture is not AI straight into a database.

Document
Extract
High confidence: continue
Review needed: human
Human-in-the-loop is not a failure of automation. It is often part of good automation.

The goal is to automate what can be automated while creating explicit handling for the cases that should not proceed on their own.

From one workflow to a reusable data layer

Once this pattern exists, the same architecture supports many workflows. In banking, KYC documents become a customer record, bank statements become a financial profile, pay stubs become income data, loan documents become application data. In healthcare, referrals become referral records, insurance documents become coverage information, lab reports become structured results, patient forms become patient records.

The input changes. The schema changes. The downstream system changes. The pattern stays the same.

Ingest, understand, structure, validate, deliver.

Why we are building this into Xpectrum

AI applications increasingly need to interact with the messy data businesses already have. Not everything lives behind a clean API. Not everything exists in a database. A huge amount of operational context is still trapped inside documents, images, PDFs, forms and text.

At the same time, extracting information is rarely the whole job. The information needs to go somewhere. Something needs to happen next.

That is why parameter extraction is part of the Xpectrum workflow layer. You accept unstructured inputs, define the structure you need, extract parameters, transform them with custom logic, and pass the results into databases, APIs, tools or subsequent workflow steps. We wrote about exposing those workflows to agents in From Any Database or API to an MCP Server in Minutes.

The model handles interpretation. The workflow handles everything after it.

The interface between AI and software is becoming structured

There is a broader shift here. The first generation of LLM applications looked like text in, text out. Then we added retrieval: a question, a lookup, an answer.

Many production AI systems now look different.

Unstructured input
AI interpretation
Structured state
Tools, APIs, databases
Real-world action
The output of AI does not always need to be language.

Sometimes the most valuable output is simply a record with a customer id, a document type, an amount, a date, a status and a next action. Because once information becomes structured, the rest of the software stack can work with it.

Unstructured in. Structured out. Action taken.

Businesses already have enormous amounts of valuable information. The problem is that much of it exists in formats software cannot directly use. AI gives us a new interpretation layer between those two worlds.

But interpretation alone is not enough. The production architecture has to turn that understanding into predictable structure, validate it, connect it to business logic, and deliver it to the systems where work happens.

PDF, image, document, text
Understand
Extract
Structure
Validate
Execute
Database, API, tool
Any supported unstructured input. Exactly the structure your application needs. Wired straight into the systems that act on it.

The future of document AI is not just about teaching models to read more documents. It is about making the information inside those documents usable by software.

XPECTRUM / ENGINEERINGMore from Xpectrum ↗