Every PDF that comes in is an hour of the team’s time
Supplier invoices that someone types into the ERP. Contracts that someone reads to extract dates and terms. Delivery notes that someone compares with orders. Technical data sheets that someone uploads to the PIM. Your team is doing human OCR.
Document IA processes those PDFs automatically — with multimodal vision, not 90s OCR. It extracts structured data that plugs into your systems. What used to take 10 minutes per document now takes 10 seconds.
How we built it
1. Audit of typical documents
You send us real samples: 20-50 documents of the type you want to process. We identify critical fields, edge cases, format variations. Before touching any code, we understand what you have and what you want to extract.
2. Extraction pipeline
We use vision-first models (Claude Vision, GPT-4 Vision) with structured output. The model reads the document as a person would and returns JSON with the fields you requested. It does not rely on prior OCR, nor on rigid templates.
3. Integration with your system
The JSON goes to your ERP, your CRM, your database, your PIM or wherever you need it. Via API or via n8n. No intermediate screens where someone copies from one place to another.
4. Human validation where it matters
Critical documents (large invoices, legal contracts) go through human review before closing. The system shows you the confidence level per field: if the model is unsure about the VAT, it tells you. It does not hide its uncertainty.
Typical document types
- Supplier invoices (any format, any language).
- Contracts with extraction of key dates, amounts, clauses.
- Product data sheets (specifications, ratios, certifications).
- Delivery notes and dispatch notes with automatic matching against orders.
- CVs with competency extraction for HR.
- Scanned tickets, receipts, supporting documents for finance.
What changes when this works
- Document processing from hours to minutes.
- Transcription errors tend to zero.
- Your administrative team is freed up for real tasks.
- Audits and compliance become traceable: every extraction is logged.
When we do NOT recommend this
- If your documents are each unique with no patterns (handwritten notes on napkins): the model cannot learn to generalise.
- If your volume is very low (<20 documents/month): the setup is not worth it.
- If your documents have legal restrictions on third-party processing (high data classification): processing them via API in the US may not be lawful.
Privacy
If your compliance policy does not allow sending documents to Claude/GPT, we use on-prem vision models (Qwen2.5-VL, InternVL3) on your infrastructure or ours in Europe. More expensive to operate, but the data never leaves.
Stack we use
- Claude Vision (Claude 4.5) for documents in general.
- GPT-4.1 Vision as an alternative in some cases.
- Qwen2.5-VL 72B on-prem for strict compliance.
- n8n + API de tu ERP/CRM for integration.
We start with a diagnostic session
Before quoting the full setup we run a 90-minute session. We look together at your real document volume, your typical formats and your destination system, and we leave with an honest recommendation: whether this service fits your current situation, or whether it makes more sense to start with something else.
We don’t charge for that session. If you’re interested, get in touch.