Your AI assistant gives wrong answers. Find the stage that breaks it.

For RAG systems and agents on your documents. Most of the time the fault is not in the language model but before it: in how the documents are read in and cut up. Windtunnel tests every stage on your data and shows what to change.

For systems in pilot or in production. Fixed prices.

Real EMA document · page 58Parser: PyMuPDF4LLM
✓ In the PDF✗ What the parser wrote
0.31
0.31
0.65
0.65
HAE Attacks
HAE Akttacs
Shape check100 %✓ every row has all its cells
Content check10th✗ last of 10 parsers

Can you answer these?

One stage at a time. One change at a time.

  1. 1

    Scrape

    What gets in

  2. 2

    Parse

    Tables, columns, scans

  3. 3

    Clean

    What gets removed

  4. 4

    Chunk

    Where text is cut

  5. 5

    Embed

    Which model, what cost

  6. 6

    Retrieve

    Fused, reranked, graded

One stage, measured: parsing

11 parsers on 22 table-heavy EMA documents. Each parser gets two grades:

  • True copy: do the tables, headings and reading order match the page?
  • Usable by an AI: with only this text, can an AI find a value and tell which row and column it belongs to?

Then what it costs and how long it takes.

ParserTrue copyUsable by an AI€ / 1,000 pagesSeconds / doc
Mistral OCR82%73%€3.431.7 s
GPT-5.4 vision + text layer77%78%€9.62not measured
GPT-5.4 vision70%77%€8.53not measured
Cohere Parse 563%not graded yet€1.295.5 s
GPT-5.4 mini vision58%65%€2.567.9 s
Azure AI Document Intelligence46%39%€8.575.3 s
pypdf43%40%€0, runs locally0.2 s
anydoc40%48%€0, runs locally0.1 s
Docling36%38%€0, your own GPU25.8 s
PyMuPDF4LLM19%19%€0, runs locally2.1 s
pdfplumber15%23%€0, runs locally0.5 s

Share of blind head-to-head comparisons won, judged by gpt-5.6-terra (5,359 judgments over all four page types). Cost at list prices in euros, measured on this set. Time per document, measured when the parser actually ran.

Clean up before you cut. Small steps that save answers.

For search, every document is cut into small pieces of text. Before that, we clean it up. Three examples from a client project with regulated technical documents:

Page headers and page numbers end up as tiny pieces of their own in the index.

Search finds empty snippets instead of the passage with the answer.

Step: remove page headers and page numbers

Pieces with no content, only a header or page number · fewer is better61

The product name at the top of every page is taken for a heading.

Pieces of text sit under the wrong heading. The AI puts a fact in the wrong section.

Step: stop treating the product name as a heading

Pieces under the wrong heading, out of all pieces · fewer is better4 of 350 of 33

Section headings get lost when the document is read in.

The AI cannot tell which section a fact comes from, so it cannot cite it.

Step: rebuild headings from the official numbering

Headings found in ten documents · more is better149178

Each step is small. Together they decide whether the AI finds the right passage. Windtunnel measures each step on its own, so you see which one is worth it.

Small samples: one document for the first two examples, ten documents of the same type for the third. They show what each check catches, not an average. Client and documents not named.

Benchmarks make the shortlist. Your data makes the choice.

A public benchmark tests somebody else's documents. Only a test on yours shows whether a model works on yours.

  1. 40models and parsers on the market
  2. Public benchmarkother documents, mostly English
  3. 4on the shortlist
  4. Windtunnelyour documents, your questions, your language
  5. 1that fits your documents

The result is not one score but, per stage: how often the right passage is found, what it costs, and how fast it is.

You keep

The corpus report

What your documents look like, and the error rate per document type, with examples.

The gold set

Labelled questions with their source passages. It is yours.

The test bench

Re-run it on every change, also as a CI gate.

The audit report

Method, data and results. A basis for your evidence to QA, notified bodies and under the EU AI Act.

You run it, we measure and improve it. Runs in your infrastructure, in the EU. No lock-in.

Fixed prices, in three steps

Step 1 · 2 weeks

Ingestion check

€9,500

Do your documents reach the system intact?

  • Error rate per document type, with examples
  • We only need read access
Fully credited against the audit
Step 2 · 5 weeks

Quality audit

€25,000

How good are your assistant's answers, and what helps most?

  • Test questions with proven answers, and they are yours
  • Your system measured against alternatives
  • Fix plan and audit report for QA and notified bodies
Includes step 1
Step 3 · after the audit

Implementation

from €35,000

We make the changes.

  • Hit rate before and after, against a target agreed up front
  • The test bench keeps running on every change
Fixed price after the audit

✓ Fits: system in pilot or in production · pharma, medical devices, energy, finance, insurance

✗ Does not fit: a first demo · you want us to run the system

Tim Treskatis

Tim Treskatis

Windtunnel, Berlin. I build and evaluate AI assistants and agents on company knowledge, also under MDR requirements. With hard documents: tables, scans and regulatory filings.

About 8 years in data and ML, 5 of them in enterprise consulting, for clients such as Airbus, Lufthansa, Deutsche Bahn and Hapag-Lloyd.

The name comes from my aerospace degree at KTH Stockholm: there, you test in a wind tunnel before you fly.

Ask for a 30-minute call

No slides. No demo. Five questions on your system, and an honest answer on fit.

Is your system in production?

Used only to reply. No newsletter. No tracking. Privacy