vllm.cpp · throughput vs the reference engine
Measured against
what each workload actually runs on
Throughput relative to the reference. 1.00 is parity, bars run from it. Higher is faster.
github.com/mudler/vllm.cpp
GB10 unless noted · greedy, reference in its own production config · docs/BENCHMARKS.md