2026-07-26
Measure vLLM 0.26.0 tuning before production rollout
Separate prefill from decode, validate execution and topology, and accept vLLM tuning only with production service evidence.
This article was source reviewed on 2026-07-26 for vLLM 0.26.0. No GPU, vLLM server, or benchmark was run for this publication. Every command below is an operator template, not a reported result.
High GPU utilization does not say whether requests are waiting, prefilling, decoding, being preempted, failing, or producing worse answers. Start from a pinned, production-shaped baseline. Change one behavior at a time. Then accept a candidate only when latency, queueing, memory, failure, quality, topology, and security all pass.
Pin the system before measuring it
Keep an evidence record for both baseline and candidate. It needs the vLLM version, immutable container digest or environment lock, model ID and commit, code and tokenizer revisions, quantization, dtype, chat template, sampling parameters, seed, and maximum model length. Record GPU model and order, driver, CUDA runtime, interconnect map, CPU and memory shape, and node count.
Also freeze tensor, pipeline, and data parallel sizes with their expected workers and GPUs. Record prefix caching, speculative configuration, KV cache allocation, max_num_seqs, batched-token limits, and client concurrency. --gpu-memory-utilization is an executor allocation fraction, not observed GPU utilization; --max-num-seqs is a per-data-parallel-rank engine-iteration limit, not client admission concurrency. Changing data-parallel size therefore changes aggregate sequence capacity. Collect queue and capacity evidence per rank and in aggregate. Those are different controls documented in the v0.26.0 engine arguments and the immutable v0.26.0 data-parallel deployment source.
The record must include the trust_remote_code setting, reviewed code commit, service identity, network boundary, cache ownership, and exposed endpoints. Keep secrets out of arguments, logs, and saved benchmark metadata. Replace every angle-bracket placeholder before use.
Operator template. Not executed for this publication.
vllm --version
nvidia-smi --query-gpu=index,name,uuid,driver_version,memory.total --format=csv,noheader
nvidia-smi topo -m
printf '%s\n' '<IMAGE_DIGEST>' '<MODEL_ID>' '<MODEL_COMMIT>' '<GPU_LAYOUT_RECORD>'Use revision pins as part of the baseline. A mutable model name is not enough to reproduce an experiment.
Operator template. Not executed for this publication.
vllm serve <MODEL_ID> \
--revision <MODEL_COMMIT> \
--code-revision <CODE_COMMIT> \
--tokenizer-revision <TOKENIZER_COMMIT> \
--dtype <DTYPE> \
--max-model-len <MAX_MODEL_LEN> \
--tensor-parallel-size <TP_SIZE> \
--pipeline-parallel-size <PP_SIZE> \
--data-parallel-size <DP_SIZE> \
--gpu-memory-utilization <GPU_MEMORY_FRACTION> \
--max-num-seqs <MAX_NUM_SEQS> \
--enable-prefix-cachingDo not add --trust-remote-code as a convenient performance variable. The v0.26.0 engine arguments default it to false. Transformers warns that custom model code executes locally and recommends reviewing and pinning it.
Shape a workload that can expose the cause
Use a sanitized production-derived request set when available. Preserve input and output lengths, prefix reuse, endpoint, sampling, and arrival distributions. A convenient random default does not establish production behavior.
| Cell | Prompt and output shape | What it isolates |
|---|---|---|
| Shared prefix, short output | Long repeated prefix with a short completion | Prefix reuse and prefill savings |
| Unique prefix, short output | Comparable input length without reusable prefixes | Prefill without cache reuse |
| Short input, long output | Minimal prefill with a long completion | Decode and inter-token behavior |
| Mixed production trace | Sanitized production length and reuse distribution | End-to-end rollout relevance |
Hold artifacts, sampling, prompt order, and client location fixed in every cell. Warm the server according to the declared production lifecycle, then exclude warmup requests. Predeclare repetitions and report every run. Sweep realistic offered load, but model the upstream admission limit separately. The v0.26.0 vllm bench serve reference defines --request-rate as request initiation and --max-concurrency as a concurrency cap; the cap can lower achieved rate when the server cannot keep up.
For a sanitized production-derived file, use --dataset-name custom with --dataset-path <SANITIZED_DATASET_PATH>. Each JSONL record needs prompt and output_tokens fields. In vLLM 0.26.0, --custom-output-len defaults to 256 and overrides per-row output lengths, so add --custom-output-len -1 to preserve the production-derived output-length distribution. Keep synthetic diagnostic choices such as random or prefix_repetition separate from that production-derived run. The accepted dataset choices, custom fields, and output-length handling are shown in the immutable v0.26.0 benchmark dataset source.
Do not use openai-chat with CustomDataset to approximate production chat requests. CustomDataset applies a chat template by default, then openai-chat wraps that resulting string as a user message before server chat templating, so the generic backend template can distort production chat requests. The custom-dataset and request behavior are in the immutable v0.26.0 benchmark dataset source and immutable v0.26.0 request-function source.
Declare TTFT, TPOT or ITL, end-to-end latency, failure, and task-quality thresholds before the run. The quality evaluator needs a fixed prompt set, scorer or reviewer protocol, sampling policy, tolerance, and treatment of stochastic variation. Save detailed JSON and run metadata so results stay tied to the full configuration. Do not compare a candidate against a baseline that used a different chat template, request body, endpoint behavior, or quality rubric.
Operator template. Not executed for this publication.
vllm bench serve \
--backend vllm \
--base-url <SERVER_BASE_URL> \
--endpoint /v1/completions \
--model <MODEL_ID> \
--dataset-name custom \
--dataset-path <SANITIZED_DATASET_PATH> \
--custom-output-len -1 \
--request-rate <TARGET_RPS> \
--max-concurrency <ADMISSION_LIMIT> \
--num-prompts <MEASURED_REQUEST_COUNT> \
--percentile-metrics ttft,tpot,itl,e2el \
--metric-percentiles 50,95,99 \
--goodput ttft:<TTFT_SLO_MS> tpot:<TPOT_SLO_MS> e2el:<E2EL_SLO_MS> \
--save-detailed \
--save-result \
--result-dir <RESULT_DIRECTORY> \
--metadata image_digest=<IMAGE_DIGEST> model_commit=<MODEL_COMMIT> config_id=<CONFIG_ID>The vLLM 0.26.0 built-in harness streams this Completions path. Production Chat Completions or non-streaming traffic requires a separate harness that reproduces the exact messages, chat template, body, headers, parameters, and stream mode. The streaming path and custom-dataset behavior are in the immutable v0.26.0 request-function source and immutable v0.26.0 benchmark dataset source.
Read prefill, queue, and decode separately
Automatic prefix caching skips reusable prompt computation. It affects prefill, not generation of new tokens, so shared-prefix results need a unique-prefix control. The v0.26.0 prefix-caching documentation describes that boundary.
| Question | Evidence |
|---|---|
| Are requests waiting? | vllm:request_queue_time_seconds, vllm:num_requests_waiting, and separately vllm:num_requests_waiting_by_reason{reason="capacity"} or {reason="deferred"} |
| Did prefill work fall? | vllm:request_prefill_time_seconds, vllm:request_prefill_kv_computed_tokens, prompt-token counters, and prefix-cache hits divided by queries |
| Did first token improve for the right reason? | vllm:time_to_first_token_seconds beside queue and prefill evidence |
| Did decode pacing improve? | vllm:request_decode_time_seconds, vllm:inter_token_latency_seconds, and benchmark TPOT or ITL |
| Is memory safe? | vllm:kv_cache_usage_perc, device observations, preemptions, OOMs, restarts, and headroom |
| Did useful service improve? | Request and token throughput constrained by TTFT, TPOT, and end-to-end goodput SLOs |
| Did reliability hold? | Benchmark error records, HTTP errors, exported vllm:request_success_total and, when enabled, vllm:corrupted_requests_total, OOMs, restarts, and timeouts |
| Did speculative work pay off? | Exported vllm:spec_decode_num_accepted_tokens_per_pos_total beside ITL, queue, memory, and goodput |
These metric names and phase meanings come from the v0.26.0 production metrics reference. Its counter names are documentation families; Prometheus exposes their counter series with _total, including vllm:request_success_total, vllm:corrupted_requests_total, and speculative-decode counters such as vllm:spec_decode_num_accepted_tokens_total and vllm:spec_decode_num_accepted_tokens_per_pos_total. Prefix-cache hit and query counters count tokens, so their ratio is a cached-token ratio, not a count of requests with a user-visible win. KV cache usage is not total physical GPU memory. Get physical headroom and OOM evidence from the device and deployment platform. A high percentage is not an acceptance criterion. The counter definitions and waiting-reason labels are in the immutable v0.26.0 metrics source and counter definitions.
vllm:corrupted_requests_total exists only when VLLM_COMPUTE_NANS_IN_LOGITS=1; that telemetry is off by default and can add compute overhead. The default and conditional metric registration are in the immutable v0.26.0 environment source and metric source.
Use the same time window and labels on both candidates. Counter rates and percentile windows must be comparable.
Operator template. Not executed for this publication.
histogram_quantile(0.95, sum by (le) (rate(vllm:time_to_first_token_seconds_bucket{model_name="<MODEL_ID>"}[<WINDOW>])))Operator template. Not executed for this publication.
sum(rate(vllm:prefix_cache_hits_total{model_name="<MODEL_ID>"}[<WINDOW>]))
/
sum(rate(vllm:prefix_cache_queries_total{model_name="<MODEL_ID>"}[<WINDOW>]))Make each change falsifiable
| Change | Hypothesis | Required confirming evidence | Reject when |
|---|---|---|---|
| Alter prefix caching | Reused prompts reduce prefill work | Cache-hit ratio rises, computed prefill tokens fall, and prefill or TTFT improves on shared-prefix traffic | Unique-prefix or decode-heavy traffic is used for a general claim, or memory, queue, preemption, failure, or quality worsens |
| Raise concurrency or scheduler capacity | More useful work completes inside the SLO | Goodput improves at the target arrival pattern without worse tail TTFT, queue, ITL, failure, or memory | Throughput rises only because requests wait longer or fail more often |
| Raise KV allocation | More KV capacity reduces eviction or preemption with headroom | KV, device memory, preemption, queue, TTFT, OOM, and restart evidence remain acceptable | Memory is merely driven closer to exhaustion |
| Change parallelism | Model placement and communication fit actual hardware | Ranks and GPUs match, collective health holds, and service metrics improve | Rank mapping, count, inter-node protection, or collective health is wrong |
| Enable speculative decoding | Accepted draft tokens reduce decode pacing for this workload | ITL or TPOT and SLO-constrained goodput improve, with acceptance, memory, failure, and quality evidence | Only utilization or raw throughput improves, acceptance is poor, or a gate regresses |
| Enable remote code | No performance hypothesis applies | Separate security approval, review, provenance, isolation, least privilege, and trusted caches | Repository code execution is treated as an ordinary tuning flag |
Speculative decoding targets inter-token latency in medium to low QPS, memory-bound workloads. It aims for losslessness, but hardware numerics, batch size, and log-probability instability can produce variation. Measure task quality on the actual workload.
Put security and topology outside the trade
If remote code is required, pin --revision, --code-revision, and --tokenizer-revision to reviewed immutable commits where applicable. Record provenance, review and scan outcomes, runtime identity, filesystem and network restrictions, and trusted cache ownership. An unreviewed execution path vetoes a performance candidate.
Verify that tensor, pipeline, and data parallel products match allocated GPUs and rank layout. Place tensor-parallel groups on the fastest available local interconnect, or justify pipeline parallelism when the hardware lacks those links. Confirm startup, rank membership, collective health, node and GPU identity, and endpoint health before benchmarking. The v0.26.0 parallelism guide covers the fit, node-boundary, and NVLink considerations.
Treat inter-node tensor, pipeline, data, PyTorch distributed, and KV-transfer traffic as insecure by default. Isolate it at the network boundary. vLLM's security guidance also documents API-key limits and unauthenticated endpoints on the same server. Expose only required routes through a reverse proxy with authentication, authorization, rate limits, and an endpoint allowlist. Match the benchmark request path, chat template, headers, parameters, and streaming mode to production as described by the OpenAI-compatible server documentation.
Roll out with an evidence gate
- Freeze the baseline manifest and acceptance thresholds.
- Run the diagnostic cells on production-equivalent hardware, changing one variable.
- Reject any candidate that fails topology or security validation.
- Compare TTFT, queue, prefill, ITL or TPOT, memory, failure, and task quality at the target arrival pattern.
- Canary on a small, explicitly bounded traffic slice with the same dashboards and error budget.
- Roll back on any predeclared stop condition. Expand only after the full observation window passes.
- Save configuration, benchmark JSON, metrics window, quality report, security review, topology record, and rollout decision together.
| Gate | Accept only when |
|---|---|
| TTFT | Target percentiles meet the predeclared SLO and do not hide extra queue time |
| Queue | Tail queue time and waiting requests remain within the declared budget at target offered load |
| Decode | ITL or TPOT meets target without a conflicting TTFT or goodput regression |
| Memory | KV and device memory retain declared headroom with no new OOM, preemption, or restart pattern |
| Failure | HTTP, benchmark, timeout, corruption, abort, and process failure rates stay within budget |
| Quality | Fixed task evaluation meets its predeclared score or tolerance under its declared stochastic method |
| Topology | Hardware, ranks, interconnect placement, collectives, and health match the design |
| Security | Provenance, pins, isolation, cache trust, endpoint exposure, and inter-node controls pass |
The gates are conjunctive. Reject the candidate when any one fails.