Technology
Four models, one learning loop, zero templates
DocuMind is built on document-native AI: models that understand layout, language and visual structure together — and a feedback loop that gets sharper with every correction your team makes.
Architecture
Document-native AI, not general AI bolted on
General-purpose LLMs hallucinate figures and lose track of tables. DocuMind's models are trained only on documents — 180 million pages of invoices, contracts, forms and reports — so they read like specialists, and every extracted value is traceable to a pixel location on the page.
- No generative step in the extraction path — values are read, never invented.
- Confidence scoring on every token, field and classification decision.
- Deterministic outputs: the same document always produces the same result.
The Model Family
Specialist models, orchestrated
Each document flows through all four models; a lightweight orchestrator decides what to run in parallel and when to escalate to human review.
LayoutLM-DM
Layout understanding
A vision-language transformer trained on 180 million document pages. Understands the relationship between text, position and visual structure — tables, headers, stamps, handwriting regions.
ReadNet
Character recognition
Our recognition backbone: printed text, cursive handwriting and degraded inputs (fax, photo, carbon copy). Emits word-level confidence and bounding boxes for every token.
ExtractQA
Field extraction
A question-answering model that treats extraction as "what does the document say about X?" — which is why it works on layouts it has never seen, with zero templates.
ClassiNet
Classification
Hierarchical classification across 81 pre-trained types plus customer-defined types learned from as few as 20 labelled examples, with full probability distributions.
The learning loop
- 1
Extract — models read and extract with confidence scores attached to every value.
- 2
Review — low-confidence fields land in your team's review queue, presented with the source region highlighted.
- 3
Learn — every correction becomes a training example for your tenant's adapted models. Nothing leaves your tenancy.
- 4
Improve — review volumes fall month over month. Typical customers reach 90%+ straight-through processing within two quarters.
Performance
Built for production volume
DocuMind runs live workloads for finance teams, hospital trusts and logistics operators — environments where a slow or unavailable pipeline stops the business. The platform is engineered accordingly.
1.8s
Median extraction latency (single invoice)
210M+
Pages processed per year
99.95%
Platform uptime, trailing 12 months
12k
Documents per minute at peak
Security & Trust
Documents are sensitive. The platform behaves like it.
Encryption everywhere
TLS 1.3 in transit, AES-256 at rest, per-tenant key isolation. Keys held in UK HSMs.
PII redaction
Automatic detection and redaction of personal data on request, with redaction events logged.
Data residency
All processing and storage in UK data centres (London & Manchester). No cross-border transfer.
Certifications
ISO 27001 certified, Cyber Essentials Plus, SOC 2 Type II report available under NDA.
Deployment
Run DocuMind where your data needs it to be
DocuMind Cloud
Our managed UK-hosted platform. Live in days, no infrastructure to run.
Private deployment
Single-tenant instances in your cloud tenancy — Azure UK South, AWS eu-west-2.
On-premises
Air-gapped appliance for the most sensitive environments. Full model parity.
Talk to our engineering team
Architecture reviews, security questionnaires, proof-of-value on your documents — our team in York handles all of it directly.