Simulation run on this device. Scenario values are not sent anywhere — the engine is fully local.
inference simulation · free
Speculative Decoding Race
Race draft-and-verify decoding against plain decoding across acceptance rates.
How often must the draft model agree for speculation to win?
Engine 1.0.0 · simulation
4
%
1095
5
5200
- E[tokens / verify]
- 2.77
- ms / committed tok
- 21.6
- Speedup vs AR
- ×1.85
- γ
- 4
Speculative decoding wins under this accept rate.
γ=4 α=0.7
Method
- Deterministic local model of the mechanism — knobs recompute outputs instantly with no network.
- Numbers are teaching approximations, not production capacity guarantees.