Shareable technical brief · August 2026

H100 prefill throughput and board-energy comparison.

Internal matched H100 prefill comparison - internal evidence, prefill only, on a same-GPU matched comparison. Independent reproduction pending. Version tags are internal milestone lineage, not products.

Summary

An internal matched comparison on a matched prefill-heavy workload measured the locked Flash, device-resident FP16 path at approximately ~1.18x the tested vLLM throughput at approximately ~12% lower board joules per prefill position, at batch 16. Throughput variation was below 1% across the measured runs. A separate internal absolute prefill champion (dual-GEMM at batch 72) measured approximately ~44,860 prefill positions per second at approximately ~0.01532 board joules per position, with no matched vLLM arm; that absolute result and this matched ratio are never combined, and the independent TESTfort baseline is a separate evaluation. This is internal evidence, prefill only, on a same-GPU matched comparison: not an independent result and not a decode or generation claim.

QuestionAnswer
Matched throughput (batch 16)~1.18x the tested vLLM throughput
Matched board energy~12% lower board J/position than the tested vLLM reference
Absolute champion (batch 72; separate; no vLLM arm)~44,860 positions per second at ~0.01532 board J/position
Run stabilityThroughput variation below 1% across measured runs
Independent reproductionPending

Peer stack used greedy sampling (temperature = 0), not a separate “deterministic product mode.” Token accounting is matched prefill positions.

Data packs: prefill_freeze_matched_20260807T210749Z (matched comparison) · prefill_accel_lock_20260807T233111Z (absolute champion; separate)

Protocol (what was compared)

SettingValue
Model classQwen2-7B-Instruct (FP16)
Sequence length128
Batch size16
WorkloadPrefill-heavy (positions = iterations × batch × 128)
HardwareSingle NVIDIA H100 80GB-class GPU
ComparisonSequential arms (not concurrent on the same GPU)
MetricsThroughput (positions/s), board energy (NVML joules / position), dual-run determinism

Peer stack: vLLM (current release in test), tensor parallel 1, prefix cache off, same token accounting.

Luxi path: production energy/throughput configuration (Flash-class attention control + device-resident multi-layer stack + FP16 weight residency). Not the high-fidelity AUDIT receipt lane.

Measured result (internal matched comparison)

Batchvs vLLM throughputvs vLLM board energy
16~1.18x the tested vLLM~12% lower board J/position

Batch-32 supporting data appears in the matched pack. No matched B72 vLLM arm yet: the batch-16 ratio is never applied to the batch-72 absolute champion. Throughput variation was below 1% across the measured runs. Board power was measured under load via standard GPU power/energy interfaces (NVML). Energy is board joules, not facility wall-plug AC.

Recipe and shape are fixed for this result: locked Flash, device-resident FP16 path; Qwen2-7B full-stack prefill; sequence length 128; batch 16. This result is never blended with the earlier independent TESTfort baseline, which is a separate evaluation of a different build.

How to read this

  • Throughput counts prompt positions under a fixed sequence length and batch (prefill-oriented), not chat tokens/s under a full OpenAI-compatible server.
  • J/position is board energy per useful position under that same definition.
  • Determinism is a pass/fail contract - dual-run agreement on the measured path (same inputs → same reported behavior).
  • This is a matched, same-day head-to-head on one GPU class - not a claim against every serving recipe, every model, or every sequence length.

Why it matters

For operators, the useful question is rarely “who wins a microkernel.” It is:

  1. Do we move more useful work per second at commercial batch?
  2. Do we spend fewer joules per unit of that work?
  3. Can we still stand behind reproducible behavior?
This matched comparison is the current internal comparison result. Independent reproduction is the next step, after which an independent claim can be made.

Scope of this result

  • Internal measurement; independent reproduction is the next step.
  • Prefill stage only; decode and generation are measured separately.
  • A matched-configuration comparison, distinct from the engine's absolute maximum performance.
  • Single-GPU matched arms, distinct from a full multi-tenant serving stack comparison (scheduling, continuous batching product surface, multi-GPU TP/PP).
  • Energy is GPU-board energy via NVML, distinct from wall-plug or PUE-adjusted energy.
  • The energy path and the high-fidelity audit path remain separate product lanes with separate evidence.
  • Chat quality versus the peer model implementation is outside this measurement's scope.

Contact

LuxiEdge · e@ewaller.com
For deeper diligence packs (methods, multi-run series, broader shapes), contact for a controlled evaluation.