Calculations run on this device. Scenario values are not sent anywhere — the engine is fully local.
inference calculator · free
LLM Serving Architect
Turn measured throughput numbers into a fleet layout for a serving stack.
Prefill, decode, or memory — what binds this deployment first?
Engine 1.0.0 · calculation
1200
10500
162048
164
- Capacity QPS
- 1.25
- Demand QPS
- 20
- Headroom
- 6%
- Verdict
- scale out
Add replicas or raise tok/s (quant, batching).
4×80 tps / 256 tok
Method
- Deterministic local model of the mechanism — knobs recompute outputs instantly with no network.
- Numbers are teaching approximations, not production capacity guarantees.