Skip to main content

Calculations run on this device. Scenario values are not sent anywhere — the engine is fully local.

research calculator · free

Bias Slice Explorer

Split an eval by attribute slices and compare deltas like an auditor.

Does accuracy hold across the slices you care about?

Engine 1.0.0 · calculation
5 pt

Slices whose accuracy delta exceeds ±budget get flagged.

Baked 1000-item eval with attribute tags. Tighten the budget to see which slices fail first.

Short prompts (<100 chars)
89.3% (+5.2)
Long prompts (>1000 chars)
69.4% (-14.7)
English
88.7% (+4.6)
Spanish
80.0% (-4.1)
Hindi
73.8% (-10.3)
Code-switching
74.3% (-9.8)
Overall accuracy
84.1%
Worst slice
Long prompts (>1000 chars)
Slices over budget
4

The average hides the failure: one slice can be 10+ points worse while the headline number looks fine. Slice first, then judge.

Method

  • Computes per-slice accuracy against the overall mean on a baked tagged eval set; slices beyond a ±5pt budget get flagged.
  • The aver­age hides failures — this is why evals ship with slice deltas, not just headline numbers.