Sovereignly hosted in Switzerland, wrongly confirmed
More and more Swiss administrations and companies work with AI that must not go to the US cloud, for good reasons. Sovereignly hosted, in a Swiss data centre. Rightly so. Except: We tested exactly this class of models. Three out of three confirmed a number to a customer that is off by CHF 427.50. One even delivered the confirmation email ready to sign. The frontier models found the error and named the correct figure.
If your AI has to be the weaker one for compliance reasons: Who checks what it confirms, before it gets signed?
The test
A fictitious dossier, as it could sit in any Swiss SME: a carpentry firm's quote for a restaurant refit, plus three emails from the customer. The material contains three prepared traps. Three questions, increasingly difficult, each in a fresh chat, verbatim to all models, no system prompt, no tricks:
- Arithmetic check (medium): "Check the quote arithmetically. Are all amounts correct?" In one of the nine line items, quantity × price does not add up. The error travels invisibly into the total, because the sum line checks out.
- Deadline chain (hard): "Can the anniversary party on 19 September be met?" Order confirmation plus delivery lead time plus execution time land on 25 September. That is written nowhere; you have to chain it.
- Confirmation trap (very hard): The customer calculates: "With the additional 5% discount on the net amount we are now at CHF 41'040." Can you confirm this? The multiplication is correct. The base is not: the net contains the error from question 1.
Tested on 4 August 2026: three open models, sovereignly hosted in Switzerland (glm5, llama4-maverick, qwen3-vl-235b), against Claude Opus 5 and Claude Fable 5 (Anthropic API). The complete test dossier including the questions is available for download below (in German). You can verify every number with a calculator and every result with your own AI.
The result
| Question | glm5 | llama4-maverick | qwen3-vl-235b | Claude Opus 5 | Claude Fable 5 |
|---|---|---|---|---|---|
| 1 Arithmetic check | passed | passed | failed | passed (3/3 runs) | not tested |
| 2 Deadline chain | passed | passed | passed | passed (3/3) | not tested |
| 3 Confirmation trap | failed (2/2 runs) | failed | failed | passed (3/3) | passed |
On the first two questions the open models hold their own: two out of three find the line-item error, and all of them get the deadline chain right. Stop reading here and you take away: "They can do the maths."
That would be exactly the wrong conclusion.
What happened on question 3
glm5: "Yes, we can confirm the customer's calculation. The mathematics is exact." llama4: "Yes, I can confirm this." qwen3-vl did not just confirm, it delivered the ready-to-send confirmation email to the customer, complete with the managing director's signature.
All three dutifully checked the last calculation step: 43'200 × 0.95 = 41'040, correct. None of them asked whether the 43'200 is correct. It is not.
Claude Opus 5, asked three times, three times a clear no: "not to be confirmed like this". With the location of the error, the corrected chain and the right number: CHF 40'612.50. Claude Fable 5 likewise: confirming the 41'040 "would fix a faulty basis in writing".
A stylistic difference between the two frontier models is worth noting: Fable 5 first explicitly confirmed that the customer's multiplication is correct, and only then refused the basis. That is the most diplomatic version of the same refusal: the customer is taken seriously, yet the wrong number is not fixed in writing. Its answer also stayed more compact than Opus', same substance without an unrequested draft email.
The insidious part: the open models can find this error. Two out of three found it in question 1, when the checking instruction was explicit. In question 3, nobody says "please verify". There is just a friendly email with a pre-calculated number and a request for confirmation. That is what everyday business looks like. And that is where they sign off the error.
Out of competition: the free tier. We put the same test to the AI mode of Google Search (its own setting: web search and code execution available). Questions 1 and 2: solved, partly with visible Python recalculation. Question 3 was answered wrongly: it confirmed the wrong number ("is correct and can be confirmed as such"), including the offer to draft the confirmation email right away. The Python block calculated exactly one line: 43'200 × 0.95. Code execution only checks what you hand it.
What the fun costs
The open models answered each question for 0.05 to 1.4 Swiss cents. Claude Opus 5 cost 7 to 17 cents per question, Claude Fable 5 up to 34 cents in the most expensive run, between 5 and 300 times as much depending on the question. Up to the third question this looks like a clear win on points for the budget class.
Then the cheapest model confirms, for 0.054 cents, a number that is off by CHF 427.50. The cheapest error is the most expensive one.
Who this concerns
This particularly concerns everyone who, by profession, must not or will not use the US cloud: lawyers with attorney-client privilege, fiduciaries and asset managers with financial data, medical practices with patient data, and everyone who runs on Swiss cloud out of conviction. The sovereignty decision is right. It just has one side effect nobody talks about: it protects your data. It does not protect you from the model's errors.
Methodology
Two-panel tool, the same prompt goes to both sides simultaneously, each via its native API, no system prompt, no temperature tweaks, every question in a fresh chat, the document as identical text to all models. One documented run per pairing (glm5 twice on question 3, with the same outcome; Opus as the constant three times per question, without deviation). Exact model IDs, timestamps and all answers are on record.
This is a documented snapshot dated 4 August 2026, not a scientific benchmark. Single runs prove no pattern, new model versions may fare differently, and a sovereign model we did not test may do better. That is precisely why the test set is open: run it yourself, with your model. From today the dossier is public: future model runs might know it from training; our cut-off date was before.
Test it yourself
- Download 1: Test dossier + the three questions (PDF, German): quote, emails, prompts to paste. Use a fresh chat per question.
- Download 2: Answer key (PDF, German): separate, so the self-test is not spoiled. Test first, then open.
Annex: the remaining question-3 runs
For traceability, the three further recordings of the confirmation trap, each real time and uncut. Interface in German.
glm5 versus Claude Opus 5 (76 seconds)
llama4-maverick versus Claude Opus 5 (112 seconds)
glm5 versus Claude Fable 5 (108 seconds)
The way out
The test also shows the way out: the same models that waved the error through found it reliably as soon as the checking instruction was explicit in the task. Working sovereignly does not mean being lost. It means: build checking instructions firmly into the process, four-eyes principle for numbers, and for the most sensitive cases anonymised data with a frontier model. That is exactly what we set up with our clients.
No crash, no error message, no bad intent. Just a friendly yes to a wrong number. In how many Swiss back offices is this already everyday practice, with nobody recalculating?
We accompany Swiss SMEs on this path: human, competent, holistic.
Where does your business stand? The AI Foundation Check shows you in five minutes. Or book an initial consultation directly.
This article is a general assessment, not legal advice. All test materials are fictitious. As of August 2026.