Gemma 4 12B RayServe GPU benchmark
RayServe and vLLM on Amazon EKS
Gemma 4 12B GPU benchmark
Six deployment configurations compared across cold and warm latency, prefill, fixed decode, and concurrency from 1 to 16.
Summary
Gemma 4 12B was deployed with Ray Serve and vLLM on Amazon EKS. The benchmark compared BF16 baselines, FP8 KV cache with n-gram speculative decoding, W4A16 quantization-aware-trained weights, and four-GPU tensor parallelism. Every configuration ran against the same prompt corpus and request schedule.
The benchmark answers two practical questions: which configuration delivers the best latency and throughput, and which lower-cost configuration is worth taking to an application quality test.
Recommendation
H100 TP=4 BASE
Use when the latency target or aggregate throughput justifies four H100 GPUs.
H100 TP=1 BASE
Strong long-prompt prefill and saturated throughput without allocating four GPUs.
L40S QAT W4A16
The fastest L40S configuration. Run an application quality gate before adoption.
Configuration and results
BASE uses BF16 weights, prefix caching, and no speculative decoding. OPT adds FP8 KV cache and five-token n-gram speculation. QAT uses W4A16 quantization-aware-trained weights with FP8 KV cache. TP is the number of GPUs serving one model replica.
Lower is better for TTFT and end-to-end latency. Higher is better for token throughput and C16 throughput per allocated hourly cost. Latency and throughput values in this comparison average the six prompt-level medians; the per-test tables below pool all raw requests.
| Rank and configuration | GPU allocation | Cold TTFT | Cold E2E | Decode | Prefill | C16 output | C16 tok/s per allocated $/hr | Allocated compute cost |
|---|---|---|---|---|---|---|---|---|
| 1 · H100 BASE TP=4 | 4x H100 | 382 ms | 3.82 s | 148.3 tok/s | 46,623 tok/s | 1,067 tok/s | 61.7 | $17.304/hr4/8 of P5 |
| 2 · H100 BASE TP=1 | 1x H100 | 896 ms | 7.23 s | 80.6 tok/s | 20,509 tok/s | 512 tok/s | 118.2 | $4.326/hr1/8 of P5 |
| 3 · H100 OPT TP=1 | 1x H100 | 1,815 ms | 7.65 s | 93.3 tok/s | 12,397 tok/s | 507 tok/s | 117.3 | $4.326/hr1/8 of P5 |
| 4 · L40S QAT W4A16 | 1x L40S | 2,871 ms | 11.15 s | 61.1 tok/s | 6,474 tok/s | 252 tok/s | 112.6 | $2.2421/hrFull G6E |
| 5 · L40S OPT | 1x L40S | 2,984 ms | 14.92 s | 44.9 tok/s | 6,265 tok/s | 227 tok/s | 101.3 | $2.2421/hrFull G6E |
| 6 · L40S BASE BF16 | 1x L40S | 3,553 ms | 22.86 s | 26.5 tok/s | 5,597 tok/s | 147 tok/s | 65.7 | $2.2421/hrFull G6E |
Cost basis: us-west-2 Linux On-Demand pricing for g6e.2xlarge and AWS Capacity Block pricing for p5.48xlarge. P5 allocation assumes all eight GPUs perform useful work; otherwise a deployment must absorb a larger share, or all, of the $34.608 hourly instance cost. Storage, data transfer, EKS, and observability are excluded. Revalidate rates before making a purchasing decision.
Concurrency-16 throughput
Average aggregate output tokens per second across the six prompt groups:
At concurrency 16, H100 TP=4 delivered 2.09 times the output throughput of H100 BASE TP=1 and 4.23 times that of L40S QAT. H100 OPT did not beat H100 BASE at concurrency 4, 8, or 16, so speculation is not a default choice for a saturated service.
How to read the metrics
TTFT
Time from request submission to the first response token. It captures prompt processing and queueing. Lower is better.
End-to-end latency
Total time from request submission until the complete response returns. Lower is better.
Prefill
Input-token processing rate before generation. It matters most for long documents and shared context. Higher is better.
Decode
Output-token generation rate for one active request after the first token. Higher is better.
C16 output
Combined output throughput from 16 simultaneous requests. It measures service capacity, not one request's speed.
P50, P90, and P95
P50 is the median. P90 and P95 expose tail behavior: 90% or 95% of requests completed at or below that value.
Complete timings by test type
Each table is calculated from all raw requests across the six prompt groups. Mean is the arithmetic average; P50 is the median; P90 and P95 show tail latency. These percentiles pool prompts of different lengths, so the tail includes the longest prompts.
T03 · Cold-prefix latency — concurrency 1
A unique prefix prevents cache reuse. Responses use natural EOS with a 512-token maximum. Lower is better.
| Deployment | Requests | Mean TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E |
|---|---|---|---|---|---|---|---|
| L40S BASE | 60 | 3,553 ms | 6,657 ms | 22.86 s | 23.53 s | 26.62 s | 26.64 s |
| L40S OPT | 60 | 2,981 ms | 5,467 ms | 14.96 s | 14.35 s | 22.78 s | 23.43 s |
| L40S QAT W4A16 | 60 | 2,872 ms | 5,225 ms | 11.17 s | 11.85 s | 14.08 s | 14.10 s |
| H100 BASE TP=1 | 60 | 898 ms | 1,627 ms | 7.21 s | 7.32 s | 8.04 s | 8.04 s |
| H100 OPT TP=1 | 60 | 1,816 ms | 3,591 ms | 7.70 s | 7.45 s | 12.00 s | 12.23 s |
| H100 BASE TP=4 | 60 | 381 ms | 636 ms | 3.81 s | 3.84 s | 4.13 s | 4.14 s |
T04 · Warm-prefix latency — concurrency 1
Each prompt is primed before measurement to evaluate reusable-prefix behavior. Responses use natural EOS with a 512-token maximum. Lower is better.
| Deployment | Requests | Mean TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E |
|---|---|---|---|---|---|---|---|
| L40S BASE | 60 | 140 ms | 172 ms | 19.44 s | 19.79 s | 20.12 s | 20.12 s |
| L40S OPT | 60 | 139 ms | 173 ms | 11.78 s | 11.04 s | 18.14 s | 18.15 s |
| L40S QAT W4A16 | 60 | 109 ms | 150 ms | 8.37 s | 8.73 s | 8.98 s | 8.98 s |
| H100 BASE TP=1 | 60 | 99 ms | 119 ms | 6.42 s | 6.46 s | 6.52 s | 6.52 s |
| H100 OPT TP=1 | 60 | 104 ms | 136 ms | 5.95 s | 5.71 s | 8.69 s | 8.70 s |
| H100 BASE TP=4 | 60 | 98 ms | 119 ms | 3.52 s | 3.53 s | 3.60 s | 3.63 s |
T05 · Fixed-length decode — concurrency 1
Every request generates exactly 512 output tokens with EOS ignored. Decode throughput excludes TTFT. Lower E2E and higher decode throughput are better.
| Deployment | Requests | Mean E2E | P50 E2E | P90 E2E | P95 E2E | P50 decode | P95 decode |
|---|---|---|---|---|---|---|---|
| L40S BASE | 60 | 22.85 s | 23.51 s | 26.55 s | 26.58 s | 26.1 tok/s | 27.9 tok/s |
| L40S OPT | 60 | 15.06 s | 14.46 s | 22.73 s | 23.36 s | 46.0 tok/s | 65.3 tok/s |
| L40S QAT W4A16 | 60 | 11.26 s | 11.84 s | 14.11 s | 14.12 s | 59.3 tok/s | 66.6 tok/s |
| H100 BASE TP=1 | 60 | 7.24 s | 7.38 s | 7.98 s | 7.99 s | 80.6 tok/s | 81.5 tok/s |
| H100 OPT TP=1 | 60 | 7.72 s | 7.45 s | 12.13 s | 12.22 s | 91.9 tok/s | 151.8 tok/s |
| H100 BASE TP=4 | 60 | 3.82 s | 3.87 s | 4.13 s | 4.15 s | 149.0 tok/s | 151.3 tok/s |
T06 · Prefill — concurrency 1
Only one output token is requested, isolating input-prompt processing. Lower TTFT and higher prefill throughput are better.
| Deployment | Requests | Mean TTFT | P50 TTFT | P90 TTFT | P95 TTFT | P50 prefill | P95 prefill |
|---|---|---|---|---|---|---|---|
| L40S BASE | 60 | 3,534 ms | 3,896 ms | 6,578 ms | 6,611 ms | 5,275 tok/s | 6,852 tok/s |
| L40S OPT | 60 | 3,016 ms | 3,383 ms | 5,488 ms | 5,503 ms | 6,074 tok/s | 7,214 tok/s |
| L40S QAT W4A16 | 60 | 2,908 ms | 3,254 ms | 5,296 ms | 5,317 ms | 6,316 tok/s | 7,228 tok/s |
| H100 BASE TP=1 | 60 | 886 ms | 988 ms | 1,611 ms | 1,612 ms | 20,781 tok/s | 21,466 tok/s |
| H100 OPT TP=1 | 60 | 1,797 ms | 1,845 ms | 3,574 ms | 3,579 ms | 11,140 tok/s | 17,169 tok/s |
| H100 BASE TP=4 | 60 | 359 ms | 401 ms | 634 ms | 636 ms | 49,518 tok/s | 51,286 tok/s |
T07 · Concurrency 1 — fixed 512-token output
Throughput reference with six measured requests per prompt. Aggregate output is averaged across the six prompt groups.
| Deployment | Requests | P50 TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E | Aggregate output |
|---|---|---|---|---|---|---|---|---|
| L40S BASE | 36 | 3,890 ms | 6,657 ms | 22.88 s | 23.55 s | 26.63 s | 26.64 s | 22.8 tok/s |
| L40S OPT | 36 | 3,351 ms | 5,512 ms | 15.20 s | 14.29 s | 23.31 s | 23.42 s | 37.0 tok/s |
| L40S QAT W4A16 | 36 | 3,235 ms | 5,300 ms | 11.26 s | 11.85 s | 14.13 s | 14.14 s | 47.9 tok/s |
| H100 BASE TP=1 | 36 | 1,019 ms | 1,627 ms | 7.22 s | 7.33 s | 7.98 s | 7.99 s | 71.4 tok/s |
| H100 OPT TP=1 | 36 | 1,874 ms | 3,589 ms | 7.71 s | 7.40 s | 11.99 s | 12.04 s | 78.2 tok/s |
| H100 BASE TP=4 | 36 | 417 ms | 634 ms | 3.79 s | 3.83 s | 4.10 s | 4.10 s | 135.5 tok/s |
T08 · Concurrency 4 — fixed 512-token output
Four simultaneous requests and twelve measured requests per prompt begin to expose queueing and shared-execution latency.
| Deployment | Requests | P50 TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E | Aggregate output |
|---|---|---|---|---|---|---|---|---|
| L40S BASE | 72 | 5,750 ms | 12,231 ms | 34.05 s | 35.32 s | 46.26 s | 52.67 s | 66.3 tok/s |
| L40S OPT | 72 | 3,676 ms | 11,215 ms | 24.65 s | 24.45 s | 40.11 s | 43.46 s | 104.5 tok/s |
| L40S QAT W4A16 | 72 | 5,518 ms | 9,871 ms | 20.04 s | 21.47 s | 29.87 s | 34.17 s | 127.4 tok/s |
| H100 BASE TP=1 | 72 | 1,976 ms | 4,151 ms | 10.03 s | 10.41 s | 12.90 s | 12.95 s | 217.4 tok/s |
| H100 OPT TP=1 | 72 | 3,248 ms | 7,083 ms | 15.56 s | 14.94 s | 28.72 s | 31.20 s | 206.0 tok/s |
| H100 BASE TP=4 | 72 | 771 ms | 1,742 ms | 4.87 s | 5.03 s | 6.01 s | 6.02 s | 438.2 tok/s |
T09 · Concurrency 8 — fixed 512-token output
Eight simultaneous requests and sixteen measured requests per prompt make tail latency increasingly important.
| Deployment | Requests | P50 TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E | Aggregate output |
|---|---|---|---|---|---|---|---|---|
| L40S BASE | 96 | 6,055 ms | 28,685 ms | 49.05 s | 45.24 s | 84.59 s | 98.71 s | 102.2 tok/s |
| L40S OPT | 96 | 5,450 ms | 24,574 ms | 36.78 s | 34.58 s | 66.25 s | 80.15 s | 159.6 tok/s |
| L40S QAT W4A16 | 96 | 5,279 ms | 23,821 ms | 32.22 s | 29.49 s | 60.16 s | 71.93 s | 185.8 tok/s |
| H100 BASE TP=1 | 96 | 2,972 ms | 6,980 ms | 13.79 s | 14.36 s | 20.45 s | 23.90 s | 344.2 tok/s |
| H100 OPT TP=1 | 96 | 3,794 ms | 14,441 ms | 23.71 s | 22.70 s | 48.35 s | 55.22 s | 335.1 tok/s |
| H100 BASE TP=4 | 96 | 1,195 ms | 2,589 ms | 6.39 s | 6.74 s | 8.69 s | 9.86 s | 709.2 tok/s |
T10 · Concurrency 16 — fixed 512-token output
Sixteen simultaneous requests and 32 measured requests per prompt represent the highest tested load. These results must not be represented by concurrency-1 E2E latency.
| Deployment | Requests | P50 TTFT | P90 TTFT | Mean E2E | P50 E2E | P90 E2E | P95 E2E | Aggregate output |
|---|---|---|---|---|---|---|---|---|
| L40S BASE | 192 | 6,324 ms | 55,024 ms | 79.18 s | 67.80 s | 155.91 s | 186.27 s | 147.3 tok/s |
| L40S OPT | 192 | 6,299 ms | 47,232 ms | 60.57 s | 53.38 s | 123.69 s | 149.76 s | 227.2 tok/s |
| L40S QAT W4A16 | 192 | 5,658 ms | 45,699 ms | 56.74 s | 47.15 s | 118.92 s | 143.81 s | 252.4 tok/s |
| H100 BASE TP=1 | 192 | 3,013 ms | 13,394 ms | 21.26 s | 19.96 s | 38.26 s | 45.06 s | 511.5 tok/s |
| H100 OPT TP=1 | 192 | 4,175 ms | 27,206 ms | 38.67 s | 34.19 s | 86.14 s | 103.44 s | 507.5 tok/s |
| H100 BASE TP=4 | 192 | 1,387 ms | 4,951 ms | 9.43 s | 9.39 s | 15.40 s | 18.06 s | 1,067.2 tok/s |
What changed the result
Performance and GPU allocation
- H100 TP=4 was the clear technical leader, but used four GPUs.
- H100 BASE TP=1 was the better one-GPU configuration for long cold prompts.
- L40S QAT led the L40S configurations and used a 9.6 GiB checkpoint.
- P5 economics assume the unallocated GPUs host other useful work.
Engine settings
- Prefix caching reduced warm TTFT to approximately 98–139 ms.
- N-gram speculation helped low-concurrency decode but hurt cold prefill.
- FP8 KV cache should be checked against application quality thresholds.
- TP=4 scaling was valuable but not linear or free.
Instance-level deployment choice
- P5: use TP=4 for the latency-sensitive replica, then place another isolated workload on the remaining GPUs. Benchmark the combined load because host, network, and storage paths are shared.
- G6E: use QAT W4A16 only after it passes a natural-EOS quality test. If it does not, L40S OPT retains BF16 weights and remains faster than the baseline.
The 512-token latency runs often reached the output cap before producing complete JSON. The performance comparison is valid, but it is not a natural-completion quality evaluation. Compare QAT with BF16 for schema validity, fact precision and recall, omissions, and hallucinations before production selection.
Method and evidence
Each deployment completed eight tests: cold-prefix latency, warm-prefix latency, fixed 512-token decode, one-token prefill, and fixed-decode concurrency at 1, 4, 8, and 16. Warm tests primed each prompt independently; cold tests used a unique nonce. The accepted artifacts contained all six prompt summaries, expected request counts, and no errors.
The TP=4 run showed four vLLM ranks on GPUs 0–3 and approximately 62.64 GiB of KV-cache capacity per rank. Host-level evidence showed those GPUs at 81–86% utilization during load while GPUs 4–7 remained unused. The QAT run confirmed the Marlin W4A16 kernel, a 72.57-second compile, and 29.68 GiB of available KV cache.
Cost interpretation
The cost-efficiency score divides concurrency-16 output throughput by the
hourly compute cost allocated to the tested GPUs. The allocation is valid only
when the remaining P5 GPUs run productive workloads. A partially idle
p5.48xlarge must assign a larger fraction of its full hourly cost to each
active model deployment. Replace these reference rates with current regional or
contracted rates before making a production decision.
Prompt corpus
The six prompts model customer-memory extraction and multi-record summarization without publishing the underlying transcripts. Four prompts perform profile-fact extraction over short (approximately 2,300 input tokens) and long (approximately 20,500 input tokens) multi-turn conversations. Each size has two strict-JSON schema variants: one records the source turn for every fact and one does not. The two hydrated prompts use longer multi-session histories: one extracts new atomic facts while deduplicating against prior learning (approximately 31,400 input tokens), and the other produces one event-focused summary with confidence, domain, and validity metadata (approximately 29,500 input tokens).
Every prompt contains task constraints, conversation context, and an explicit output schema; hydrated cases also exercise prior-learning merge or deduplication behavior. Every deployment used the same prompt files and request schedule. The files remain local because the transcript content is not needed to interpret the benchmark results.
Run the benchmark
Start with the Gemma example README to prepare, deploy, and verify a scenario. Then follow the benchmark runbook to execute T03–T10, validate each result artifact, collect the results, and clean up the deployment.
Benchmark run: August 2026.