A model can describe a successful action without leaving the application in the right state. We tested a concrete workflow: read a document, call a posting tool, and verify the row actually saved in SQLite.
Both models failed our frozen zero-failure policy. That result needs context: some failures were wrong amounts or missing posts; others were formatting mismatches that do not establish a wrong monetary value. The study also found bugs in our evaluation integration. Those distinctions are the useful result.
Build on existing work
We reused CORD v2, the public receipt dataset, and pdfhell's existing invoice generators. Hugging Face Datasets handled revision-pinned loading. Inspect handled model requests, tools, native logs, offline scoring and retry. Multivon connected that evidence to explicit required checks and acceptance decisions.
There is no new dataset engine or agent scheduler here. CORD remains a receipt parsing dataset; our ledger task is a separate projection of its labels, with its own limitations. Dataset attribution and changes are recorded in the reproduction bundle.
What we ran
The held-out selection contained 19 CORD receipts and 20 synthetic invoices. One preselected receipt had no scalar total label and was excluded without replacement. Each source had two input treatments, evaluated by Haiku 4.5 and Sonnet 5: 156 scored cases from 39 source documents.
For receipts, we compared resized images with the dataset's annotated words. That text baseline loses layout and unannotated text; it is not a full OCR system. For invoices, we compared native PDFs with locally rendered 150 DPI images. Each model had one generation and a maximum of 512 output tokens per case.
| Persisted outcome check | Haiku 4.5 | Sonnet 5 |
|---|---|---|
| Receipt image | 16/19 | 18/19 |
| Receipt annotated text | 18/19 | 17/19 |
| Currency conversion, PDF | 9/10 | 10/10 |
| Currency conversion, pixels | 9/10 | 10/10 |
| Hidden OCR conflict, PDF | 10/10 | 10/10 |
| Hidden OCR conflict, pixels | 10/10 | 10/10 |
These samples are small. For example, 10/10 has a Wilson 95% interval of 72.2–100%; 18/19 has an interval of 75.4–99.1%. The full analysis reports every interval and paired comparison. It supports neither a general model ranking nor production readiness. New synthetic seeds share the same layouts, and CORD merchant independence is unknown.
A dataset label is not a business rule
One Haiku result saved USD 13,456.57 for an invoice requiring USD 13,456.78. A separate database query caught the wrong amount. A Sonnet response exhausted its output budget, omitted a required tool field and saved nothing.
Other failures were different. An upstream total label was Rp. 91,000; both
models' text-track outputs saved 91,000. Another response changed a separator
while preserving the digits. Those fail our verbatim contract, but should not
all be presented as wrong payment amounts.
That exposes a weakness in our task mapping. Industrial use needs explicit currency, locale, rounding and normalization rules. Borrow a credible dataset, then check whether its labels answer your actual acceptance question. We kept the frozen results and published a separate failure review. We did not change held-out labels to improve the scores. The visual review was performed by an AI assistant, not an independent human adjudicator.
Repair scoring without giving the model another chance
A malformed tool call exposed a bug in our Inspect bridge: it serialized tool errors using the wrong object interface. The native log preserved the response. We fixed the bridge and regraded that output offline; it remained a failure. Inspect then preserved 16 completed generations and ran 62 remaining samples. Original grading errors and interrupted attempts remain in the evidence.
Across development and held-out runs, 188 model events had recorded usage, estimated at $0.6998 at API list prices. Two cancelled requests have unknown billing. This is not an invoice total, and preserved requests were counted once across retry logs.
Where this leaves Multivon
With a correct posting tool, checking its arguments and checking the final row agreed on these outcomes. This study does not prove an advantage over a good argument checker. A separate crash/retry experiment tested a duplicate-write fault, where persisted-state verification matters.
Our direction is to make task evidence useful for release decisions: reuse established datasets and runtimes, verify side effects, expose domain assumptions, and preserve failures when measurement breaks. A maintained collection of independently reviewed workflow assertions could become valuable. A small sandbox study does not prove a moat or customer demand.
Case manifests, acceptance policies and Inspect integration are now available in multivon-eval 0.18.0 on PyPI. The study preserves its original development revision in the execution lock. Start with the frozen protocol and runnable study and the document workflow guide.
The raw evidence bundle and offline reproduction recipe include the source assets, native logs, ledger state and checksums. Recompute the reported analysis without model calls, or inspect the original trajectories with Inspect's existing viewer. Dataset attribution and the study's limitations travel with the archive.