A statistically rigorous, causal evaluation layer for LLM apps, built on top of DeepEval.
Created and maintained by @routsom.
DeepEval measures. It gives you a score. But a single score can't tell you whether it's real (or just noise), why your app produced an output, or whether the LLM judge that produced the score can be trusted.
causeval wraps DeepEval's metrics - they stay the measurement instrument - and adds the three things a score alone can't give you:
- 📊 Uncertainty - Is the score real? Repeated sampling, variance decomposition, clustered-bootstrap confidence intervals, paired comparisons, and statistically valid gates.
- 🔬 Causality - What caused this output or failure? RAG context ablation and counterfactual context, input perturbations with metamorphic relations, and agent step-level counterfactual replay.
- ⚖️ Judge validity - Can we trust the judge? Bias audits, calibration against human labels, prediction-powered inference (PPI), and conformal abstention.
The one rule that drives everything: causeval never reports a bare score. Every public result carries an item count, a repeat count, a confidence interval, the method used, and full provenance (versions, seed, dataset hash, git SHA).
- Why causeval?
- Install
- Quickstart
- What it adds, feature by feature
- Does it actually work? (validation benchmarks)
- CLI
- Design principles
- Project status & roadmap
- Author
- License
DeepEval is an excellent measurement library. causeval is not a competitor - it is the layer that turns DeepEval's measurements into decisions you can defend in a code review, a launch meeting, or a paper. Here is precisely what it adds:
causeval targets Python 3.10-3.12.
# from source (recommended while in alpha)
git clone https://github.com/routsom/causeval.git
cd causeval
uv sync # core install
uv sync --all-extras # + optional backends (see below)Optional extras, each pulling in a heavier dependency only where you need it:
The core estimators (bootstrap CIs, PPI, IRT, AIPW) are implemented natively on numpy/scipy - the extras are for optional cross-checks and NLI, not for the math.
The core object is Experiment: run each metric over each item R times and get back a
RunResult whose estimates carry CIs and variance components.
from causeval import Experiment, MetricSpec
exp = Experiment(
dataset=goldens, # DeepEval Goldens or plain dicts
metrics=[MetricSpec("FaithfulnessMetric", {"threshold": 0.7})],
repeats=5, # resample the judge 5x per item
seed=0,
)
run = exp.run() # or: await exp.a_run()
print(run.summary())Run 'run' (seed=0)
judge=gpt-4o-mini causeval=0.0.1
Faithfulness [base]: 0.812 [95% CI 0.771, 0.849] (n=50, R=5, icc=0.34, flaky=6)
method=cluster_bootstrap_studentized_B=2000
You immediately learn what a bare score hides: the CI, that ~34% of the variance is between items (the rest is judge noise), and that 6 items are flaky (pass rate between 0.2 and 0.8).
from causeval.stats import compare, gate
comparisons = compare( # Holm-adjusted across metrics
baseline_run.measurements,
candidate_run.measurements,
margin=0.02,
)
report = gate(comparisons)
print(report.overall) # "pass" | "regression" | "inconclusive"
raise SystemExit(report.exit_code) # 0 / 1 / 2 for CIThe gate is three-valued on purpose: "inconclusive" (the CI straddles the margin) is a
first-class outcome, so you never ship a regression that hid behind noise, and never block a
release on a difference you didn't have the power to detect. Plan that power up front with
causeval.stats.plan_power.
DeepEval's Faithfulness tells you the answer is entailed by the context. It cannot tell you the model would have said the same thing without it. causeval intervenes on the context and watches the answer move:
from causeval.interventions import ground, GroundingItem
result = ground(my_rag_app, dataset, repeats=8)
print(result.summary())RAG grounding (tau=0.5, policy=follow_context)
context_reliance: +0.463 [95% CI +0.311, +0.621] (n=40, R=8)
counterfactual_adherence: +0.506 [95% CI +0.372, +0.643] (n=40, R=8)
classes: grounded=20, parametric=20
Counterfactual Adherence edits one supporting fact to a plausible-but-false value and checks whether the answer follows the edit. Items are labelled grounded / parametric / confabulating / mixed, and any answer that passes Faithfulness while being parametric is flagged false-faithful.
from causeval.judge_audit import calibrate, ppi_mean_ci
report = calibrate(judge_scores, human_labels) # Spearman, kappa, isotonic map, ECE + CI
effect = ppi_mean_ci(human_labels, judge_on_labeled, judge_on_unlabeled) # human mean, debiasedPPI gives you the human-quality mean with a CI that is unbiased (unlike averaging the judge) yet far tighter than using your few human labels alone.
Every statistical claim has a benchmark with a known ground truth, run in CI on synthetic
data (fakes, no network), with a live variant for real models. These are the committed offline
numbers (bench/results/):
(B1 - CI coverage, false-regression rate, and power - is verified by simulation tests under
tests/.)
The numbers above are the offline benchmarks: synthetic fakes with a known ground truth,
run in CI on every commit so the statistics are provably correct without spending a token.
Each benchmark also has a live variant that swaps the fake for a real judge/model, so you
can publish the same claims against, say, gpt-4o-mini or claude-haiku.
🚧 Status: partial. The baseline row below is a real live run against Claude Haiku 4.5; B2-B6 live runs are still pending a larger budget (tracked in
PROGRESS.md). Run any of them yourself with the commands underneath.
¹ Two real live runs against Claude Haiku 4.5. Clean run (n=6, R=5, temp 0,
baseline_live.md): every item a confident 1.000 with zero flip - the reassuring baseline, and an end-to-end pipeline confirmation. Borderline run (n=5, R=5, temp 1.0, deliberately partial/ambiguous items,baseline_live_borderline.md) surfaces what a bare score hides: AnswerRelevancy drops to 0.833 with a wide CI [0.63, 1.00] (a single item's score is genuinely uncertain across the set), ContextualRelevancy correctly falls to 0.17 on the off-topic contexts, and one item wobbled across repeats. Two findings you can only get by measuring: within-item flip is ≈ 0 even at temperature 1, so Haiku 4.5 is a stable judge; but it rated deliberately unsupported claims (invented patent counts, a made-up budget) as fully Faithful = 1.000 - a real judge-leniency signal that causeval's judge audit (bias probes, calibration, PPI) exists to catch.² Live judge-bias probes against Claude Haiku 4.5 (n=12 neutral answer pairs judged in both orders; n=12 pointwise items;
b3_judge_bias_live.md). The headline is real and significant: a position effect of −0.292 [−0.422, −0.126] means that, on answer pairs of equal quality, Haiku picks the second-presented option ~79% of the time - a textbook LLM-judge position bias, and exactly the kind of thing you must correct for before trusting a pairwise judge. It also mildly penalizes padded/verbose answers (−0.061), and shows no significant formatting or authorship-label effect. Unlike the offline B3 (which injects a known 0.15/0.08 bias to prove the probes recover it), the live run measures whatever bias the real judge actually has.³ Live RAG grounding against a real Claude Haiku 4.5 RAG app (
b2_rag_grounding_live.md), using the SPEC's fictional-vs-well-known design. Counterfactual Adherence (does the answer follow a false edit to a supporting fact?) separates the two groups with AUROC 0.833: all 6 fictional items score CA 1.0 with high Context Reliance (the model must use the context), while the well-known items mostly resist the false edit (CA ≈ 0, CR ≈ 0 - the model already knows the answer). It lands below the offline 1.000 / the 0.9 target for an honest reason on real data: on 2 of 12 items the model was swayed by the false context (e.g. it accepted "Romeo and Juliet was written by Dickens") - a real sycophancy signal the metric surfaces. DeepEval Faithfulness would rate every edited-context answer "faithful" and could not make this grounded-vs-parametric distinction at all.⁴ Live PPI (
b4_ppi_live.md) over 40 factual-QA items (20 correct, 20 with a plausible-but-wrong answer) where objective 0/1 correctness is the "human label" and Claude's pointwise score is the predictorf. From a random 12-item labeled subset, PPI estimates the true mean correctness as 0.510 [0.353, 0.667] (truth = 0.500) with an effective sample size ≈ 41 - i.e. 12 human labels bought the precision of ~41, and the CI is half the width of the human-only estimate (0.31 vs 0.58). Honest caveat: on these clear-cut items Claude was a well-calibrated grader (naive judge-mean 0.503, essentially unbiased), so there was little bias to correct here - unlike the offline synthetic judge. PPI's win on this run is label efficiency; its bias-correction matters most on the subtler tasks where judges drift (see the faithfulness leniency in the baseline footnote).⁵ Live IRT (
b6_irt_live.md) using 6 real Claude models as the systems (haiku-4-5, sonnet-4-5, sonnet-5, opus-4-5, opus-4-8, fable-5) on 30 hard short-answer items. Fitting 2PL and pruning to the most-informative 15 items preserved the system ranking exactly (Kendall τ = 1.000). Honest caveat: these are all frontier models, so they cluster near ceiling (accuracy 0.93-1.00) and the true ability spread is narrow - the ranking is close, so preserving it is a lighter test than the offline B6, which validates pruning across a wide simulated ability range. The live run is a real end-to-end confirmation that Fisher-information pruning doesn't scramble the ranking; the offline B6 is the rigorous one.⁶ Live agent attribution (
b5_agent_attribution_live.md) against a real Claude Haiku tool-agent (price → multiply → add-tax tasks) with a wrong tool argument injected at a known step. Counterfactual replay localized the injected fault as the decisive step in 4 of 5 valid tasks (0.80). It lands below the offline 100% for two honest, interesting reasons: (1) real Claude agents often self-correct an injected fault (one task was dropped as "no persistent failure" because the agent noticed and redid the step - a genuine robustness finding), and (2) this environment's Anthropic SDK build rejectstemperature=0, so the replay rollouts are noisy. The offline B5 (a deterministic scripted agent) validates the attribution engine at 100% over 40 tasks.
Reproduce a live run (needs an API key for the provider you name; nothing is hardcoded):
export OPENAI_API_KEY=... # or your provider's key
uv sync
# baseline flakiness with a real judge
uv run python -m causeval.bench.baseline_variance --model gpt-4o-mini
# the full live test suite (opt-in; excluded from the default offline run)
uv run pytest -m liveLive results are written under bench/results/ alongside the offline ones;
open a PR with your table and we'll add it here.
Every command writes a JSON result and prints a human-readable summary; none of them will ever print a bare number.
causeval run --config eval.yaml --out runs/ # repeated-sampling run
causeval compare --baseline a.json --candidate b.json --margin 0.02
causeval gate --baseline a.json --candidate b.json --margin 0.02 # exit 0/1/2
causeval plan --pilot-baseline a.json --pilot-candidate b.json --metric Faithfulness --detect 0.03
causeval ground --config rag.yaml --out runs/ # RAG causal grounding
causeval audit-judge --config judge.yaml --out runs/ # calibration + PPI
causeval attribute --config agent.yaml --out runs/ # agent step-level blamecauseval gate sets the process exit code (0 pass, 1 regression, 2 inconclusive), so it
drops straight into CI as a release gate.
These are enforced by tests, not just documented (see CLAUDE.md):
- Wrap DeepEval, never fork it. Depend on deepeval>=4.2,<5; never patch its internals.
- All DeepEval imports go through one guarded module that sets telemetry/dotenv opt-outs before the first import. A test enforces that no other module imports it directly.
- A fresh metric object per measurement - never share an instance across concurrent calls (DeepEval issue #3356).
- Never report a bare score. Every result carries n_items,n_repeats, a CI, the method, and provenance - including CLI output.
- No telemetry, no network, no side effects at import time.
- No hardcoded model prices or names in logic - they come from user config.
- Offline tests never call real LLMs - they use configurable fakes; live tests are opt-in.
- Every statistical method has a simulation test proving its coverage or error rate on data with a known answer.
Alpha. The full estimator suite (Phases 0-6) and the causal extensions of Phase 7 are
implemented, with 180+ offline tests and strict typing. causeval is a working name.
- ✅ Phase 0-1: scaffold, core schemas, adapters, statistics (CIs, gates, planning)
- ✅ Phase 2: RAG causal grounding
- ✅ Phase 3: judge audit (bias, calibration, PPI, conformal, jury)
- ✅ Phase 4: perturbations + deterministic checks
- ✅ Phase 5: agent step-level attribution
- ✅ Phase 6: IRT pruning + adaptive sampling
- ✅ Phase 7 (partial): CoT faithfulness, observational AIPW, OpenTelemetry trace import
- ⏳ Planned: framework harness adapters (LangGraph / OpenAI Agents / Pydantic AI), an HTML report, a pytest plugin, and published live-model benchmark tables
See SPEC.md for the full design and PROGRESS.md for the
decision log.
uv sync --all-extras
uv run pytest -m "not live" # fast offline suite (must always pass)
uv run pytest -m live # real LLM calls; needs API keys
uv run ruff check . && uv run ruff format --check .
uv run mypy src/causeval # strict
uv run python -m causeval.bench.rag_grounding # regenerate a benchmark reportContributions are welcome. See CONTRIBUTING.md for setup, the checks your
PR must pass, and the non-negotiable rules (especially: never report a bare score, and every
statistical method needs a simulation test). Use the issue templates to
report a bug or
request a feature.
causeval is created, designed, and maintained by @routsom.
If this project is useful to you, please ⭐ star the repo and follow @routsom for more work on rigorous LLM evaluation. Issues, ideas, and pull requests are welcome.
Apache 2.0, matching DeepEval. Portions of DeepEval, where copied, retain their
Apache 2.0 headers and are recorded in NOTICE.
causeval is an independent project and is not affiliated with or endorsed by Confident AI, the maintainers of DeepEval.