Skip to main content

Gemma 4 12B RayServe GPU benchmark

RayServe and vLLM on Amazon EKS

Gemma 4 12B GPU benchmark

Six deployment configurations compared across cold and warm latency, prefill, fixed decode, and concurrency from 1 to 16.

48 primary tests6 prompt groups0 request errorsL40S and H100TP=1 and TP=4

Summary

Gemma 4 12B was deployed with Ray Serve and vLLM on Amazon EKS. The benchmark compared BF16 baselines, FP8 KV cache with n-gram speculative decoding, W4A16 quantization-aware-trained weights, and four-GPU tensor parallelism. Every configuration ran against the same prompt corpus and request schedule.

The benchmark answers two practical questions: which configuration delivers the best latency and throughput, and which lower-cost configuration is worth taking to an application quality test.

Recommendation

Performance leader

H100 TP=4 BASE

Use when the latency target or aggregate throughput justifies four H100 GPUs.

1,067 tok/sConcurrency-16 aggregate output
Best single H100

H100 TP=1 BASE

Strong long-prompt prefill and saturated throughput without allocating four GPUs.

896 msAverage of prompt-level median cold TTFT
L40S candidate

L40S QAT W4A16

The fastest L40S configuration. Run an application quality gate before adoption.

61 tok/sSingle-request fixed decode

Configuration and results

BASE uses BF16 weights, prefix caching, and no speculative decoding. OPT adds FP8 KV cache and five-token n-gram speculation. QAT uses W4A16 quantization-aware-trained weights with FP8 KV cache. TP is the number of GPUs serving one model replica.

Lower is better for TTFT and end-to-end latency. Higher is better for token throughput and C16 throughput per allocated hourly cost. Latency and throughput values in this comparison average the six prompt-level medians; the per-test tables below pool all raw requests.

Rank and configurationGPU allocationCold TTFTCold E2EDecodePrefillC16 outputC16 tok/s per allocated $/hrAllocated compute cost
1 · H100 BASE TP=44x H100382 ms3.82 s148.3 tok/s46,623 tok/s1,067 tok/s61.7$17.304/hr4/8 of P5
2 · H100 BASE TP=11x H100896 ms7.23 s80.6 tok/s20,509 tok/s512 tok/s118.2$4.326/hr1/8 of P5
3 · H100 OPT TP=11x H1001,815 ms7.65 s93.3 tok/s12,397 tok/s507 tok/s117.3$4.326/hr1/8 of P5
4 · L40S QAT W4A161x L40S2,871 ms11.15 s61.1 tok/s6,474 tok/s252 tok/s112.6$2.2421/hrFull G6E
5 · L40S OPT1x L40S2,984 ms14.92 s44.9 tok/s6,265 tok/s227 tok/s101.3$2.2421/hrFull G6E
6 · L40S BASE BF161x L40S3,553 ms22.86 s26.5 tok/s5,597 tok/s147 tok/s65.7$2.2421/hrFull G6E

Cost basis: us-west-2 Linux On-Demand pricing for g6e.2xlarge and AWS Capacity Block pricing for p5.48xlarge. P5 allocation assumes all eight GPUs perform useful work; otherwise a deployment must absorb a larger share, or all, of the $34.608 hourly instance cost. Storage, data transfer, EKS, and observability are excluded. Revalidate rates before making a purchasing decision.

Concurrency-16 throughput

Average aggregate output tokens per second across the six prompt groups:

H100 TP=4 BASE
1,067
H100 TP=1 BASE
512
H100 TP=1 OPT
507
L40S QAT W4A16
252
L40S OPT
227
L40S BASE BF16
147

At concurrency 16, H100 TP=4 delivered 2.09 times the output throughput of H100 BASE TP=1 and 4.23 times that of L40S QAT. H100 OPT did not beat H100 BASE at concurrency 4, 8, or 16, so speculation is not a default choice for a saturated service.

How to read the metrics

TTFT

Time from request submission to the first response token. It captures prompt processing and queueing. Lower is better.

End-to-end latency

Total time from request submission until the complete response returns. Lower is better.

Prefill

Input-token processing rate before generation. It matters most for long documents and shared context. Higher is better.

Decode

Output-token generation rate for one active request after the first token. Higher is better.

C16 output

Combined output throughput from 16 simultaneous requests. It measures service capacity, not one request's speed.

P50, P90, and P95

P50 is the median. P90 and P95 expose tail behavior: 90% or 95% of requests completed at or below that value.

Complete timings by test type

Each table is calculated from all raw requests across the six prompt groups. Mean is the arithmetic average; P50 is the median; P90 and P95 show tail latency. These percentiles pool prompts of different lengths, so the tail includes the longest prompts.

T03 · Cold-prefix latency — concurrency 1

A unique prefix prevents cache reuse. Responses use natural EOS with a 512-token maximum. Lower is better.

DeploymentRequestsMean TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2E
L40S BASE603,553 ms6,657 ms22.86 s23.53 s26.62 s26.64 s
L40S OPT602,981 ms5,467 ms14.96 s14.35 s22.78 s23.43 s
L40S QAT W4A16602,872 ms5,225 ms11.17 s11.85 s14.08 s14.10 s
H100 BASE TP=160898 ms1,627 ms7.21 s7.32 s8.04 s8.04 s
H100 OPT TP=1601,816 ms3,591 ms7.70 s7.45 s12.00 s12.23 s
H100 BASE TP=460381 ms636 ms3.81 s3.84 s4.13 s4.14 s

T04 · Warm-prefix latency — concurrency 1

Each prompt is primed before measurement to evaluate reusable-prefix behavior. Responses use natural EOS with a 512-token maximum. Lower is better.

DeploymentRequestsMean TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2E
L40S BASE60140 ms172 ms19.44 s19.79 s20.12 s20.12 s
L40S OPT60139 ms173 ms11.78 s11.04 s18.14 s18.15 s
L40S QAT W4A1660109 ms150 ms8.37 s8.73 s8.98 s8.98 s
H100 BASE TP=16099 ms119 ms6.42 s6.46 s6.52 s6.52 s
H100 OPT TP=160104 ms136 ms5.95 s5.71 s8.69 s8.70 s
H100 BASE TP=46098 ms119 ms3.52 s3.53 s3.60 s3.63 s

T05 · Fixed-length decode — concurrency 1

Every request generates exactly 512 output tokens with EOS ignored. Decode throughput excludes TTFT. Lower E2E and higher decode throughput are better.

DeploymentRequestsMean E2EP50 E2EP90 E2EP95 E2EP50 decodeP95 decode
L40S BASE6022.85 s23.51 s26.55 s26.58 s26.1 tok/s27.9 tok/s
L40S OPT6015.06 s14.46 s22.73 s23.36 s46.0 tok/s65.3 tok/s
L40S QAT W4A166011.26 s11.84 s14.11 s14.12 s59.3 tok/s66.6 tok/s
H100 BASE TP=1607.24 s7.38 s7.98 s7.99 s80.6 tok/s81.5 tok/s
H100 OPT TP=1607.72 s7.45 s12.13 s12.22 s91.9 tok/s151.8 tok/s
H100 BASE TP=4603.82 s3.87 s4.13 s4.15 s149.0 tok/s151.3 tok/s

T06 · Prefill — concurrency 1

Only one output token is requested, isolating input-prompt processing. Lower TTFT and higher prefill throughput are better.

DeploymentRequestsMean TTFTP50 TTFTP90 TTFTP95 TTFTP50 prefillP95 prefill
L40S BASE603,534 ms3,896 ms6,578 ms6,611 ms5,275 tok/s6,852 tok/s
L40S OPT603,016 ms3,383 ms5,488 ms5,503 ms6,074 tok/s7,214 tok/s
L40S QAT W4A16602,908 ms3,254 ms5,296 ms5,317 ms6,316 tok/s7,228 tok/s
H100 BASE TP=160886 ms988 ms1,611 ms1,612 ms20,781 tok/s21,466 tok/s
H100 OPT TP=1601,797 ms1,845 ms3,574 ms3,579 ms11,140 tok/s17,169 tok/s
H100 BASE TP=460359 ms401 ms634 ms636 ms49,518 tok/s51,286 tok/s

T07 · Concurrency 1 — fixed 512-token output

Throughput reference with six measured requests per prompt. Aggregate output is averaged across the six prompt groups.

DeploymentRequestsP50 TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2EAggregate output
L40S BASE363,890 ms6,657 ms22.88 s23.55 s26.63 s26.64 s22.8 tok/s
L40S OPT363,351 ms5,512 ms15.20 s14.29 s23.31 s23.42 s37.0 tok/s
L40S QAT W4A16363,235 ms5,300 ms11.26 s11.85 s14.13 s14.14 s47.9 tok/s
H100 BASE TP=1361,019 ms1,627 ms7.22 s7.33 s7.98 s7.99 s71.4 tok/s
H100 OPT TP=1361,874 ms3,589 ms7.71 s7.40 s11.99 s12.04 s78.2 tok/s
H100 BASE TP=436417 ms634 ms3.79 s3.83 s4.10 s4.10 s135.5 tok/s

T08 · Concurrency 4 — fixed 512-token output

Four simultaneous requests and twelve measured requests per prompt begin to expose queueing and shared-execution latency.

DeploymentRequestsP50 TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2EAggregate output
L40S BASE725,750 ms12,231 ms34.05 s35.32 s46.26 s52.67 s66.3 tok/s
L40S OPT723,676 ms11,215 ms24.65 s24.45 s40.11 s43.46 s104.5 tok/s
L40S QAT W4A16725,518 ms9,871 ms20.04 s21.47 s29.87 s34.17 s127.4 tok/s
H100 BASE TP=1721,976 ms4,151 ms10.03 s10.41 s12.90 s12.95 s217.4 tok/s
H100 OPT TP=1723,248 ms7,083 ms15.56 s14.94 s28.72 s31.20 s206.0 tok/s
H100 BASE TP=472771 ms1,742 ms4.87 s5.03 s6.01 s6.02 s438.2 tok/s

T09 · Concurrency 8 — fixed 512-token output

Eight simultaneous requests and sixteen measured requests per prompt make tail latency increasingly important.

DeploymentRequestsP50 TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2EAggregate output
L40S BASE966,055 ms28,685 ms49.05 s45.24 s84.59 s98.71 s102.2 tok/s
L40S OPT965,450 ms24,574 ms36.78 s34.58 s66.25 s80.15 s159.6 tok/s
L40S QAT W4A16965,279 ms23,821 ms32.22 s29.49 s60.16 s71.93 s185.8 tok/s
H100 BASE TP=1962,972 ms6,980 ms13.79 s14.36 s20.45 s23.90 s344.2 tok/s
H100 OPT TP=1963,794 ms14,441 ms23.71 s22.70 s48.35 s55.22 s335.1 tok/s
H100 BASE TP=4961,195 ms2,589 ms6.39 s6.74 s8.69 s9.86 s709.2 tok/s

T10 · Concurrency 16 — fixed 512-token output

Sixteen simultaneous requests and 32 measured requests per prompt represent the highest tested load. These results must not be represented by concurrency-1 E2E latency.

DeploymentRequestsP50 TTFTP90 TTFTMean E2EP50 E2EP90 E2EP95 E2EAggregate output
L40S BASE1926,324 ms55,024 ms79.18 s67.80 s155.91 s186.27 s147.3 tok/s
L40S OPT1926,299 ms47,232 ms60.57 s53.38 s123.69 s149.76 s227.2 tok/s
L40S QAT W4A161925,658 ms45,699 ms56.74 s47.15 s118.92 s143.81 s252.4 tok/s
H100 BASE TP=11923,013 ms13,394 ms21.26 s19.96 s38.26 s45.06 s511.5 tok/s
H100 OPT TP=11924,175 ms27,206 ms38.67 s34.19 s86.14 s103.44 s507.5 tok/s
H100 BASE TP=41921,387 ms4,951 ms9.43 s9.39 s15.40 s18.06 s1,067.2 tok/s

What changed the result

Performance and GPU allocation

  • H100 TP=4 was the clear technical leader, but used four GPUs.
  • H100 BASE TP=1 was the better one-GPU configuration for long cold prompts.
  • L40S QAT led the L40S configurations and used a 9.6 GiB checkpoint.
  • P5 economics assume the unallocated GPUs host other useful work.

Engine settings

  • Prefix caching reduced warm TTFT to approximately 98–139 ms.
  • N-gram speculation helped low-concurrency decode but hurt cold prefill.
  • FP8 KV cache should be checked against application quality thresholds.
  • TP=4 scaling was valuable but not linear or free.

Instance-level deployment choice

  • P5: use TP=4 for the latency-sensitive replica, then place another isolated workload on the remaining GPUs. Benchmark the combined load because host, network, and storage paths are shared.
  • G6E: use QAT W4A16 only after it passes a natural-EOS quality test. If it does not, L40S OPT retains BF16 weights and remains faster than the baseline.
Quality gate required

The 512-token latency runs often reached the output cap before producing complete JSON. The performance comparison is valid, but it is not a natural-completion quality evaluation. Compare QAT with BF16 for schema validity, fact precision and recall, omissions, and hallucinations before production selection.

Method and evidence

Each deployment completed eight tests: cold-prefix latency, warm-prefix latency, fixed 512-token decode, one-token prefill, and fixed-decode concurrency at 1, 4, 8, and 16. Warm tests primed each prompt independently; cold tests used a unique nonce. The accepted artifacts contained all six prompt summaries, expected request counts, and no errors.

The TP=4 run showed four vLLM ranks on GPUs 0–3 and approximately 62.64 GiB of KV-cache capacity per rank. Host-level evidence showed those GPUs at 81–86% utilization during load while GPUs 4–7 remained unused. The QAT run confirmed the Marlin W4A16 kernel, a 72.57-second compile, and 29.68 GiB of available KV cache.

48 result artifactsExpected request countsZero request errorsIndependent prompt primingChecksum manifestsPhysical TP=4 GPU evidence

Cost interpretation

The cost-efficiency score divides concurrency-16 output throughput by the hourly compute cost allocated to the tested GPUs. The allocation is valid only when the remaining P5 GPUs run productive workloads. A partially idle p5.48xlarge must assign a larger fraction of its full hourly cost to each active model deployment. Replace these reference rates with current regional or contracted rates before making a production decision.

Prompt corpus

The six prompts model customer-memory extraction and multi-record summarization without publishing the underlying transcripts. Four prompts perform profile-fact extraction over short (approximately 2,300 input tokens) and long (approximately 20,500 input tokens) multi-turn conversations. Each size has two strict-JSON schema variants: one records the source turn for every fact and one does not. The two hydrated prompts use longer multi-session histories: one extracts new atomic facts while deduplicating against prior learning (approximately 31,400 input tokens), and the other produces one event-focused summary with confidence, domain, and validity metadata (approximately 29,500 input tokens).

Every prompt contains task constraints, conversation context, and an explicit output schema; hydrated cases also exercise prior-learning merge or deduplication behavior. Every deployment used the same prompt files and request schedule. The files remain local because the transcript content is not needed to interpret the benchmark results.

Run the benchmark

Start with the Gemma example README to prepare, deploy, and verify a scenario. Then follow the benchmark runbook to execute T03–T10, validate each result artifact, collect the results, and clean up the deployment.

Benchmark run: August 2026.