Shareable technical brief · August 2026
H100 prefill throughput and board-energy comparison.
Internal matched H100 prefill comparison - internal evidence, prefill only, on a same-GPU matched comparison. Independent reproduction pending. Version tags are internal milestone lineage, not products.
Summary
An internal matched comparison on a matched prefill-heavy workload measured the locked Flash, device-resident FP16 path at approximately ~1.18x the tested vLLM throughput at approximately ~12% lower board joules per prefill position, at batch 16. Throughput variation was below 1% across the measured runs. A separate internal absolute prefill champion (dual-GEMM at batch 72) measured approximately ~44,860 prefill positions per second at approximately ~0.01532 board joules per position, with no matched vLLM arm; that absolute result and this matched ratio are never combined, and the independent TESTfort baseline is a separate evaluation. This is internal evidence, prefill only, on a same-GPU matched comparison: not an independent result and not a decode or generation claim.
| Question | Answer |
|---|---|
| Matched throughput (batch 16) | ~1.18x the tested vLLM throughput |
| Matched board energy | ~12% lower board J/position than the tested vLLM reference |
| Absolute champion (batch 72; separate; no vLLM arm) | ~44,860 positions per second at ~0.01532 board J/position |
| Run stability | Throughput variation below 1% across measured runs |
| Independent reproduction | Pending |
Peer stack used greedy sampling (temperature = 0), not a separate
“deterministic product mode.” Token accounting is matched prefill positions.
Data packs: prefill_freeze_matched_20260807T210749Z (matched comparison) · prefill_accel_lock_20260807T233111Z (absolute champion; separate)
Protocol (what was compared)
| Setting | Value |
|---|---|
| Model class | Qwen2-7B-Instruct (FP16) |
| Sequence length | 128 |
| Batch size | 16 |
| Workload | Prefill-heavy (positions = iterations × batch × 128) |
| Hardware | Single NVIDIA H100 80GB-class GPU |
| Comparison | Sequential arms (not concurrent on the same GPU) |
| Metrics | Throughput (positions/s), board energy (NVML joules / position), dual-run determinism |
Peer stack: vLLM (current release in test), tensor parallel 1, prefix cache off, same token accounting.
Luxi path: production energy/throughput configuration (Flash-class attention control + device-resident multi-layer stack + FP16 weight residency). Not the high-fidelity AUDIT receipt lane.
Measured result (internal matched comparison)
| Batch | vs vLLM throughput | vs vLLM board energy |
|---|---|---|
| 16 | ~1.18x the tested vLLM | ~12% lower board J/position |
Batch-32 supporting data appears in the matched pack. No matched B72 vLLM arm yet: the batch-16 ratio is never applied to the batch-72 absolute champion. Throughput variation was below 1% across the measured runs. Board power was measured under load via standard GPU power/energy interfaces (NVML). Energy is board joules, not facility wall-plug AC.
Recipe and shape are fixed for this result: locked Flash, device-resident FP16 path; Qwen2-7B full-stack prefill; sequence length 128; batch 16. This result is never blended with the earlier independent TESTfort baseline, which is a separate evaluation of a different build.
How to read this
- Throughput counts prompt positions under a fixed sequence length and batch (prefill-oriented), not chat tokens/s under a full OpenAI-compatible server.
- J/position is board energy per useful position under that same definition.
- Determinism is a pass/fail contract - dual-run agreement on the measured path (same inputs → same reported behavior).
- This is a matched, same-day head-to-head on one GPU class - not a claim against every serving recipe, every model, or every sequence length.
Why it matters
For operators, the useful question is rarely “who wins a microkernel.” It is:
- Do we move more useful work per second at commercial batch?
- Do we spend fewer joules per unit of that work?
- Can we still stand behind reproducible behavior?
Scope of this result
- Internal measurement; independent reproduction is the next step.
- Prefill stage only; decode and generation are measured separately.
- A matched-configuration comparison, distinct from the engine's absolute maximum performance.
- Single-GPU matched arms, distinct from a full multi-tenant serving stack comparison (scheduling, continuous batching product surface, multi-GPU TP/PP).
- Energy is GPU-board energy via NVML, distinct from wall-plug or PUE-adjusted energy.
- The energy path and the high-fidelity audit path remain separate product lanes with separate evidence.
- Chat quality versus the peer model implementation is outside this measurement's scope.
Contact
LuxiEdge · e@ewaller.com
For deeper diligence packs (methods, multi-run series, broader shapes), contact for a
controlled evaluation.