From Unstructured Data to Structured Data: Building AI Data Pipelines for Banking and Healthcare
Turn faxes, PDFs, documents, images, forms and free-form text into structured, validated data, and send it straight into the systems that need it. Extraction is one node in a workflow, not the whole job.
Most production systems do not operate on documents. They operate on structure.
A loan processing system expects fields like applicant_name, annual_income, account_balance and loan_amount. A healthcare workflow might expect patient_name, date_of_birth, provider, diagnosis, insurance_information and referral_reason.
But the information entering these organisations rarely arrives that cleanly. It arrives inside PDFs, scanned documents, images, forms, clinical notes, bank statements, pay stubs, referral letters, insurance documents and free-form text. The information is there. The structure is not.
For years, organisations have closed that gap with some combination of OCR, document templates, regex, custom parsers, manual data entry and increasingly complex document-processing pipelines. Multimodal AI changes what is possible. But getting a model to understand a document is only the beginning.
The problem is not reading the document
Consider a bank statement. A multimodal model can look at it and explain what it contains: a statement for September, with account information, opening and closing balances, deposits, withdrawals and transaction history.
That is impressive. It is also not particularly useful to a production system. A lending workflow does not need a paragraph explaining the document. It needs fields.
{
"account_holder": "John Smith",
"account_type": "Checking",
"statement_start_date": "2026-09-01",
"statement_end_date": "2026-09-30",
"opening_balance": 38240.12,
"closing_balance": 42850.72,
"total_deposits": 12450.00,
"total_withdrawals": 7839.40,
"currency": "USD"
}That is a fundamentally different problem. We are moving from document understanding to schema-constrained extraction to doing something with the result.
And that last step matters. The structured result usually is not the final product. It needs to be validated, transformed, stored, passed to another API, reviewed by a person, or used to trigger another workflow.
Why banking and healthcare make this obvious
Almost every industry deals with unstructured information, but banking and healthcare make the challenge particularly visible. Both operate sophisticated structured systems while continuously receiving information in highly variable formats.
Banking
A lending or financial workflow may need to process bank statements, loan applications, mortgage documents, pay stubs, tax documents, KYC documents, invoices, financial statements, transaction reports and scanned forms. The data eventually needs to become fields another system can understand.
Healthcare
Healthcare organisations encounter referral documents, patient intake forms, lab reports, clinical notes, insurance documents, discharge summaries, medical records, prescriptions, scanned forms and PDFs received from external providers.
And a great deal of it still arrives by fax. Fax remains one of the most common ways referrals and records move between providers in the United States, which means a large share of inbound clinical information lands as a scanned image of a printed page: no text layer, variable quality, no structure at all. A fax inbox is an unstructured input like any other, and it is often the highest-volume one.
Again, the destination is usually structured. The source is not. That mismatch creates an enormous amount of glue code and operational work.
From a document to a schema
Take a simplified banking example. A lending application receives three files: a bank statement, a pay stub and a loan application. Different documents, different layouts, potentially different institutions.
But the lending system does not care about their layouts. It cares about a defined set of information.
{
"applicant": {
"full_name": "",
"address": ""
},
"employment": {
"employer": "",
"monthly_income": 0
},
"financials": {
"account_balance": 0,
"monthly_deposits": 0
},
"loan": {
"requested_amount": 0,
"purpose": ""
}
}The input can remain messy. The output cannot. Instead of asking “what does this document say?” we are asking “given this input, populate this exact structure.”
Building this in Xpectrum
In Xpectrum this is a workflow. Extraction is not an isolated AI call; it is one node in something that runs.

The trigger panel is worth a look. The inputs are declared as fields the workflow accepts, so a paragraph of free text, an uploaded document and an uploaded image are all first-class entry points rather than special cases.
Step 1: accept unstructured input
The workflow starts with the data the organisation already has: text, an image, a document, a PDF, or another supported input.
A banking workflow might receive a loan application, a bank statement and a pay stub. A healthcare workflow might receive a referral PDF, an insurance document and a clinical note.
The workflow does not need every source to share the same structure before processing begins. That is the point.
Step 2: define what you actually want
This is the most important step. Instead of letting the model decide what information is interesting, the developer defines what the application needs.
{
"patient": {
"name": "string",
"date_of_birth": "date"
},
"referral": {
"referring_provider": "string",
"specialty": "string",
"reason": "string",
"priority": "string"
},
"insurance": {
"provider": "string",
"member_id": "string"
}
}
The extractor takes an input variable from the trigger and a list of parameters, each one named, typed and required or not. The instruction attached to a field is where you pin down the detail that would otherwise drift between documents.
The parameter extractor maps the unstructured input into those parameters. Your database does not need to understand a PDF. Your API does not need to interpret a clinical note. Your application receives the fields it was designed to consume.
Step 3: normalise the result
Extraction is rarely the end of the pipeline. Real documents are messy. A date may arrive as 09/14/26, September 14, 2026 or 14-Sep-2026. Currency may arrive as $12,450, USD 12,450.00 or 12.45K USD. Names, addresses, identifiers and categorical values all vary.
So after extraction, the result passes through a custom function. “September 14, 2026” becomes “2026-09-14”. “$12,450.00” becomes 12450.00. One source’s “cardiac specialist” becomes the “Cardiology” the receiving system expects.
That is a better production pattern than asking the model to own every step.
Step 4: validate before anything downstream
Extraction should not automatically imply trust. Suppose a mortgage workflow extracts a monthly income, an account balance and a requested loan amount. Before that reaches another system, the workflow may need to check whether required fields are present, whether numeric values are actually numeric, whether dates are valid, whether expected identifiers exist, whether the result satisfies application-specific rules, and whether the case should go to a person.
This is where production AI starts to look different from an AI demo. The objective is not to make the model answer. It is to make the system behave predictably.
LangChain makes a similar distinction in its production-agent writing: model intelligence is one component, and production systems need the surrounding runtime, controls, observability and operational plumbing. LangChain ↗ The same principle applies here. Extraction is one capability inside a larger system.
Step 5: put the data where work happens
Once information has been extracted and transformed, it should not stop inside the AI platform. It becomes input to the next system.
Example: bank statement processing
A customer submits a bank statement. The raw input might contain customer information, account numbers, the statement period, balances, deposits, withdrawals, transactions, fees and a good deal of boilerplate.
The application may only need a fraction of it.
{
"account_holder": "John Smith",
"statement_period": {
"from": "2026-09-01",
"to": "2026-09-30"
},
"opening_balance": 38240.12,
"closing_balance": 42850.72,
"total_deposits": 12450.00,
"total_withdrawals": 7839.40
}Example: healthcare referral intake
A referral might arrive as a fax or a PDF containing patient demographics, insurance information, provider information, medical context and the reason for referral. The receiving system may only need a handful of fields.
{
"patient_name": "Jane Doe",
"date_of_birth": "1982-04-16",
"referring_provider": "Dr. Smith",
"specialty": "Cardiology",
"reason_for_referral": "Cardiac evaluation",
"insurance_provider": "Example Health",
"priority": "Routine"
}Note what the extracted record contains: a name, a date of birth, a member id. That is protected health information, which makes the next section less optional than it looks.
Security is part of the pipeline, not a wrapper around it
Document pipelines tend to carry the most sensitive material an organisation holds. A referral contains a name, a date of birth and a member id. A bank statement contains an account number and a full transaction history. A KYC packet contains identity documents.
Extraction does not reduce that exposure. It concentrates it. One workflow now reads every document, holds the fields it pulled out, and writes them into several systems. That is a single place worth getting right, and a single place worth auditing.
What the architecture has to account for
- Scoped credentials. The workflow should act with permissions that belong to the caller or to the workflow itself, not with a blanket key that can reach every record in the system.
- A narrow blast radius for the model. The extractor needs the document and the schema. It does not need standing access to the database it eventually writes to.
- An audit trail. Which document was processed, which fields came out, which system received them, and who approved the ones that needed approval.
- Retention limits. The source document, the intermediate extraction and the logs all persist somewhere. Each needs a deliberate lifetime rather than an accidental one.
- Encryption in transit and at rest, including the intermediate state between extraction and the write, which is easy to overlook because it feels transient.
- Human review as a control, not just a quality step. For some categories the right answer is that nothing reaches a system of record without a person signing it off.
For workflows touching electronic protected health information in the United States, the HIPAA Security Rule requires administrative, physical and technical safeguards for ePHI. Fax makes this concrete: the moment a referral lands in an inbound fax queue it is already PHI, before any model has looked at it, which means the queue, the storage behind it and the processing that follows all sit inside the same obligation.
The useful consequence of putting extraction inside a workflow is that these controls live at the workflow layer, where they apply to every document the same way, rather than being reimplemented per integration.
Extraction is only one node
This is perhaps the most important design decision. Unstructured-to-structured extraction should not be a standalone application. It should be a primitive inside a larger workflow.
PDF to extract to validate to API is useful. But real business processes get more interesting quickly.
One intelligence layer, deterministic logic around it
Models are valuable precisely because the inputs are not predictable. Traditional software works extremely well when field A is here, field B is there, and field C always follows a known format. Real documents do not behave that way.
AI gives us a flexible interpretation layer. But once the document has been interpreted, much of the remaining process can, and often should, be deterministic.
- Validation
- Business rules
- Transformation
- Permissions
What happens when extraction is not certain?
Production extraction should not pretend uncertainty does not exist. Documents can be incomplete. Scans can be poor. Two values can conflict. A required field may not exist. A document might not even be the type you expected.
For high-impact workflows, the right architecture is not AI straight into a database.
The goal is to automate what can be automated while creating explicit handling for the cases that should not proceed on their own.
From one workflow to a reusable data layer
Once this pattern exists, the same architecture supports many workflows. In banking, KYC documents become a customer record, bank statements become a financial profile, pay stubs become income data, loan documents become application data. In healthcare, referrals become referral records, insurance documents become coverage information, lab reports become structured results, patient forms become patient records.
The input changes. The schema changes. The downstream system changes. The pattern stays the same.
Why we are building this into Xpectrum
AI applications increasingly need to interact with the messy data businesses already have. Not everything lives behind a clean API. Not everything exists in a database. A huge amount of operational context is still trapped inside documents, images, PDFs, forms and text.
At the same time, extracting information is rarely the whole job. The information needs to go somewhere. Something needs to happen next.
That is why parameter extraction is part of the Xpectrum workflow layer. You accept unstructured inputs, define the structure you need, extract parameters, transform them with custom logic, and pass the results into databases, APIs, tools or subsequent workflow steps. We wrote about exposing those workflows to agents in From Any Database or API to an MCP Server in Minutes.
The interface between AI and software is becoming structured
There is a broader shift here. The first generation of LLM applications looked like text in, text out. Then we added retrieval: a question, a lookup, an answer.
Many production AI systems now look different.
Sometimes the most valuable output is simply a record with a customer id, a document type, an amount, a date, a status and a next action. Because once information becomes structured, the rest of the software stack can work with it.
Unstructured in. Structured out. Action taken.
Businesses already have enormous amounts of valuable information. The problem is that much of it exists in formats software cannot directly use. AI gives us a new interpretation layer between those two worlds.
But interpretation alone is not enough. The production architecture has to turn that understanding into predictable structure, validate it, connect it to business logic, and deliver it to the systems where work happens.
The future of document AI is not just about teaching models to read more documents. It is about making the information inside those documents usable by software.