Document extraction
ExtractIQDocument extraction that pulls structured data out of forms, invoices and scanned files.
ExtractIQ is a document extraction system. It reads whatever arrives — a clean PDF, a phone photograph of a delivery note, a fax from 2004 — finds the fields you care about, and writes them into your software.
There's a working version of this you can try right now — ExtractIQ
The problem
Somebody is retyping this by hand
Every business has a person, or several, whose week goes on moving numbers from documents into software. It is slow, it is dull, and it is where the mistakes come from — a transposed figure on an invoice is not noticed until it is a problem.
Templates were meant to solve this. But every supplier lays their invoice out differently, and the moment one of them redesigns theirs the template breaks and it is back to typing.
A page goes in.A row comes out.
What it does
What the document extraction system does
Reads documents never meant to be read by a machine
Scans, photographs taken at an angle, faxes, handwriting on a printed form, and tables that run across three pages.
Pulls the fields you actually need
You define what matters — invoice number, net, VAT, line items, dates, policy numbers — and get those back as clean typed values rather than a wall of text.
Knows when it is unsure
Every field carries a confidence score. Anything below the line you set goes to a person to check, instead of quietly entering your accounts wrong.
Puts the data where it belongs
Straight into your accounting system, ERP or database. The end of the process is a record in your software, not a spreadsheet somebody still has to import.
How it works
How a document becomes data
Agree the fields and the tolerance
We start from the output rather than the input — which fields, what type, what a valid value looks like, and what an error would actually cost. That last one decides where the human-review line sits.
Read the document as it arrives
Deskew, clean up, handle the multi-page and the photographed. There is no template per supplier, so a redesigned invoice does not break the pipeline the morning it turns up.
Check before it commits
Extracted values are tested against rules you already have — does the VAT add up, does the purchase order exist, is the date plausible. Failures and low-confidence fields queue for review.
Write into your system, with a trail
Every posted record links back to the page and the region it was read from, so a query or an audit can always reach the original.
Where it fits
Where the retyping happens
Accounts payable
Supplier invoices from two hundred suppliers in two hundred layouts, matched against purchase orders and posted without anyone retyping them.
Freight and customs
Bills of lading, packing lists and commercial invoices, turned into the fields a customs declaration needs.
Lending and mortgages
Payslips, bank statements and identity documents read into an application file, with the figures pulled out for affordability checks.
Clinical and laboratory administration
Referral forms and results that arrive as scans, typed into the patient record instead of onto a pile.
What you get
A pipeline, and a queue for the exceptions
A running extraction pipeline
Connected to wherever documents arrive — an inbox, a scanner, a folder, an API — and processing them without anyone having to start it.
An exception queue
One screen where a person handles only the documents the system flagged, with the page and the uncertain field side by side.
Accuracy measured on your documents
We test against a sample of your real paperwork, bad scans included, and report field-level accuracy before go-live rather than after.
Handover and a support window
Documentation, a walkthrough for the team running the queue, and a tuning period once real volume is flowing.
See it in action
ExtractIQ is on the work page — a short video of it running, and a demo you can use yourself.
Try it live — ExtractIQOften built alongside
- KnowledgeCoreEnterprise RAG
Once the paperwork is structured, KnowledgeCore makes it answerable — so what did we agree with this supplier becomes a question you can ask.
- AgentMeshMulti-agent automation
Extraction is usually one step in something longer. AgentMesh handles what happens after the data lands: matching, approving, chasing, posting.
Tell us what you are retyping.
Which documents, roughly how many a week, and what happens to the data once it is in. We will tell you what can be automated, what will always need a human eye, and where the review line ought to sit.