Evaluation evidence · updated September 17, 2026
What does the evidence
show before shipping?
We test workflow choices and evaluator judgments against public labels, then publish the methods, failures, cost and uncertainty. The current studies include a negative result: asking one response to produce both an answer and its evidence record made financial QA worse. None of these studies establishes a state-of-the-art or customer-validity claim.
What this means
Prompt structure and evaluator defaults can both change the result. Check them against task labels, preserve failures, and measure the operational cost before using a score as a release gate.
Regulated-finance target · complete released test population
Evidence in the same response reduced answer quality
September 17, 2026 · 1,663 TAT-QA questions across 277 financial contexts. Both treatments use Claude Sonnet 5 and the official TAT-QA scorer. The model sees tables, paragraphs and questions; gold answers, scales, derivations, facts and mappings are excluded.
| Treatment | Exact match | F1 | Estimated cost | Median latency |
|---|---|---|---|---|
| Answer only | 74.74% | 82.76% | $1.9648 | 4.57 s |
| Answer + evidence | 72.22% | 79.96% | $2.7841 | 6.18 s |
Joint generation lost 2.53 percentage points of exact match and 2.81 points of F1. The paired context-bootstrap 95% intervals are −4.21 to −0.84 and −4.37 to −1.31 points. It also cost 41.70% more and increased median latency 35.24%. This rejects the tested prompt design; it does not prove that a separate evidence-binding stage will improve results.
The evidence response used valid in-range locations on 99.58% of questions, but only 55.20% met the strict rule of an exact answer plus complete agreement with released support locations. Location agreement is not semantic proof, and alternate valid evidence may exist.
Public test-label contamination is unknown, the model name is a mutable provider alias, and TAT-QA supplies contexts. This is a workflow study, not a leaderboard, retrieval result or regulated-customer validation.
Inspect the protocol, numeric results and limitations →Current study: RAGChecker human agreement
September 17, 2026 · 280 comparison cases, two human annotations each. Both methods use the same pinned Haiku 4.5 model, reference answers and responses. Scores below are Pearson correlations, not accuracy percentages.
| Method | Correlation | Scored responses |
|---|---|---|
| Multivon AnswerAccuracy | 0.499 | 560 / 560 |
| Direct rating, frozen parser | 0.452 | 513 / 560 |
The paired 95% interval for the correlation difference is −0.031 to +0.129, so this run does not establish a winner. The direct baseline uses median imputation for 37 pairs with missing scores. A post-hoc leading-integer parser recovers 44 replies and raises its correlation to 0.533, reversing the point-estimate ordering; three responses remain unscored. This diagnostic does not replace the frozen result. QAG used four calls per response and direct rating used one; more calls did not establish better judging.
Intervals resample cases within domains and keep both annotations together. This is one maintainer-run configuration on public data, with unmeasured model contamination and source-document dependence. It does not validate regulated customer workflows.
Read the protocol, paired predictions and limitations →Historical pilot: framework defaults
Last run: 2026-06-26. These configurations use different prompts, thresholds and integrations. The standalone harness is not public, so this historical comparison cannot yet be independently reproduced and does not establish a general framework ranking.
Dataset: ragtruth-sum (n=100)(RAG-Truth summarization split, 100-case stratified sample with human labels). Each row is one framework's faithfulness/hallucination metric scored against the human labels at the framework's default threshold.
Judge: claude-haiku-4-5
| Framework | Threshold | F1 | Precision | Recall | Latency (ms) | Errors |
|---|---|---|---|---|---|---|
| multivon-eval | 0.90 | 0.690 | 0.615 | 0.784 | 10066 | 0 |
| DeepEval | 0.50 (default) | — | — | — | — | 100 |
| RAGAS | — | — | — | — | — | 0 |
Judge: gpt-4o-mini
| Framework | Threshold | F1 | Precision | Recall | Latency (ms) | Errors |
|---|---|---|---|---|---|---|
| multivon-eval | 0.90 | 0.729 | 0.912 | 0.608 | 10341 | 0 |
| DeepEval | 0.50 (default) | 0.038 | 1.000 | 0.020 | 12720 | 0 |
| RAGAS | 0.50 (default) | 0.038 | 1.000 | 0.020 | 136122 | 4 |
Reading the table.At default thresholds, DeepEval has no measurable F1 with the claude-haiku judge because every case errors. With gpt-4o-mini it scores F1 0.038 (recall 0.02; it flags almost none of the labeled hallucinations). multivon-eval's F1 is 0.690 (claude-haiku) and 0.729 (gpt-4o-mini). Default-vs-default is the comparison most users get when they install each framework and run with the documented configuration.
What happens if you tune the threshold?
Some of DeepEval's poor performance at default settings is a threshold issue. Below, F1 across a threshold sweep on the gpt-4o-mini judge. multivon-eval's best F1 is 0.837 at threshold 0.95. DeepEval's best F1 is 0.609 at threshold 0.95. Even at best-tuned thresholds, multivon-eval has a ~37% F1 advantage. Threshold sweeps are computed on the test set, so read them as upper bounds, not held-out estimates.
| Threshold | multivon-eval F1 | DeepEval F1 |
|---|---|---|
| 0.30 | 0.000 | 0.000 |
| 0.50 | 0.000 | 0.038 |
| 0.60 | 0.075 | 0.107 |
| 0.70 | 0.210 | 0.194 |
| 0.80 | 0.418 | 0.384 |
| 0.90 | 0.729 | 0.569 |
| 0.95 | 0.837 | 0.609 |
Inspect the recorded result
Download the exact configuration and table above as JSON. The standalone cross-framework harness is not public yet, so we do not label this run independently reproducible. The public SDK benchmark directory contains related evaluator benchmarks and raw artifacts.
curl -O https://multivon.ai/data/framework-benchmark-summary.json
# Related public SDK benchmarks and raw artifacts:
git clone https://github.com/multivon-ai/multivon-eval
cd multivon-eval/benchmarksDatasets: HaluEval QA and Summarization (100-case stratified samples each), plus the ragtruth-sum split. Judges tested: Claude Haiku 4.5, GPT-4o-mini. Same random seed across all runs. Inspect the machine-readable summary or the public SDK benchmarks.
Calls we made
- Same judge for all frameworks.We don't let each framework use a different judge model. If we did, the comparison would measure judge quality, not framework quality.
- Same threshold semantics.Each framework's documented "default threshold" was used as-is, then thresholds were swept in the second section so you can see how each behaves at its best.
- RAGAS included as of ragas 0.4.3. It errored on every case in the prior harness; it now completes (4/100 cases still error, disclosed above) and runs ~15× slower than the other two.
- Where multivon-eval loses, it's documented. See COMMENTARY.mdin the repo for cases where multivon-eval flagged a hallucination the human label said was correct, or vice versa. We don't hide them.