How we tested it: four AI models against an M&A contract file
Method and findings for the viewpoint «The insidious trap»: how we measured four AI models against a fictitious contract file.
The setup
The basis is a fictitious contract file for a company acquisition (share purchase and shareholders' agreement, management buy-out, around 15 pages). All names, companies, figures and numbers are invented; no real client, no real data set.
It is built on two levels: a data-protection level with a wealth of personal details (the test bed for anonymisation) and an analysis level with built-in arithmetic, date and contractual contradictions, some disguised so that a superficial check confirms them rather than exposes them.
Before the analysis we anonymised the file with our Sovereign Cockpit: names, accounts and identification numbers never leave the building, while the figures stay fully checkable. That way even cloud models work only on pseudonymised data.
We measured every model answer against a solution key prepared in advance.
How we measured
The same task went to four models, each through its own interface: two open-weight models of the kind you would run in your own data centre (Mistral-Small, 119 billion parameters despite the name, around 136 GB of GPU memory across two H100 GPUs; Qwen3.5, around 442 GB across four H200), and two frontier models from the cloud (Claude Opus 4.8 and Fable 5, both with adjustable reasoning depth, the «effort» we varied from standard to maximum). The brief: summarise the file as an adviser to a financial investor and sanity-check the figures, stating facts only when fully certain and otherwise flagging the uncertainty.
We recomputed every answer deterministically rather than judging it by eye, and varied two operating levers: breaking up the request and the reasoning depth.
What we found
On pure arithmetic the open-weight models partly kept up: Qwen3.5 caught the hidden calculation errors cleanly. Mistral-Small, by contrast, found none, took wrong core figures as fact and added errors of its own, in the same confident advisory tone as its correct answers. That is the deceptive confidence: an answer that sounds checked and isn't.
The frontier models covered more, above all the contractual logic: which clause contradicts another, which figure does not fit the document. But none was infallible here either. Claude Opus confirmed a disguised calculation error with a tick and carried it forward, because a wrong intermediate assumption happened to match the claimed final figure. The most thorough runs were Fable at full reasoning depth; they even found errors that had escaped us when we built the file.
Two things held for every model: repetition is no control (two runs did not converge on the truth, they produced new variants), and operation is a second lever (focused sub-questions and more reasoning depth clearly raised the hit rate for the same model).
What follows
- A frontier model more reliably spots which figure is wrong and where a contract contradicts itself; that is what the higher price buys.
- Trust no figure unchecked. The reliable sequence: compute deterministically, let the model interpret, have a human verify against the original.
- For a decision-maker, the wrong figure in a confident tone is more dangerous than an open «I can't say that for sure».
Transparency
The file is fictitious; all personal, company and payment data are invented. The model assessments are traceable against a deterministically computed target. Part of the evaluation was co-written by one of the tested models; the conflict of interest is disclosed, the figures remain independently verifiable.
The most dangerous AI answer is the one that is wrong and sounds sure.
iConference AG · Method for the viewpoint «The insidious trap» · July 2026