vllm.cpp · throughput vs the reference engine

Measured against what each workload actually runs on

Throughput relative to the reference. 1.00 is parity, bars run from it. Higher is faster.
github.com/mudler/vllm.cpp GB10 unless noted · greedy, reference in its own production config · docs/BENCHMARKS.md