Calculations run on this device. Scenario values are not sent anywhere — the engine is fully local.
research calculator · free
Bias Slice Explorer
Split an eval by attribute slices and compare deltas like an auditor.
Does accuracy hold across the slices you care about?
Engine 1.0.0 · calculation
5 pt
Slices whose accuracy delta exceeds ±budget get flagged.
Baked 1000-item eval with attribute tags. Tighten the budget to see which slices fail first.
- Overall accuracy
- 84.1%
- Worst slice
- Long prompts (>1000 chars)
- Slices over budget
- 4
The average hides the failure: one slice can be 10+ points worse while the headline number looks fine. Slice first, then judge.
Method
- Computes per-slice accuracy against the overall mean on a baked tagged eval set; slices beyond a ±5pt budget get flagged.
- The average hides failures — this is why evals ship with slice deltas, not just headline numbers.