Skip to main content

Calculations run on this device. Scenario values are not sent anywhere — the engine is fully local.

inference calculator · free

LLM Serving Architect

Turn measured throughput numbers into a fleet layout for a serving stack.

Prefill, decode, or memory — what binds this deployment first?

Engine 1.0.0 · calculation
1200
10500
162048
164
Capacity QPS
1.25
Demand QPS
20
Headroom
6%
Verdict
scale out

Add replicas or raise tok/s (quant, batching).

4×80 tps / 256 tok

Method

  • Deterministic local model of the mechanism — knobs recompute outputs instantly with no network.
  • Numbers are teaching approximations, not production capacity guarantees.