KV-cache sizing estimate

Dependency-free planning aid · keyboard accessible · responsive

Assumptions / Hypothèses

Formula / Formule: bytes ≈ 2 × layers × KV heads × head dimension × cached tokens × bytes/element × sequences

Approximation only. Real tokenizers, kernels, architectures, runtimes, cache sharing, metadata, fragmentation, and quality vary. Measure the exact target system; do not use this estimate as a capacity or quality guarantee.