Skip to main content

Simulation run on this device. Scenario values are not sent anywhere — the engine is fully local.

inference simulation · free

Speculative Decoding Race

Race draft-and-verify decoding against plain decoding across acceptance rates.

How often must the draft model agree for speculation to win?

Engine 1.0.0 · simulation
4
%
1095
5
5200
E[tokens / verify]
2.77
ms / committed tok
21.6
Speedup vs AR
×1.85
γ
4

Speculative decoding wins under this accept rate.

γ=4 α=0.7

Method

  • Deterministic local model of the mechanism — knobs recompute outputs instantly with no network.
  • Numbers are teaching approximations, not production capacity guarantees.