Three audiences · one mission
Energy-efficient AI inference with public H100 evidence.
~1.17-1.18× throughput · ~10-14% lower board J/position vs vLLM†
Batch 16 and 32 · Qwen2-7B-Instruct FP16 · H100 · Version 100
Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.
Higher throughput and lower board energy than the tested vLLM stack on this prefill workload. Faithful Qwen2-7B CUDA correctness gates passed.
†Version 100 internal measurement, July 2026. Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, one generated token per iteration. Batch 16: LuxiEdge ~41,700 pos/sec, ~0.0171 J/pos vs vLLM ~35,800 pos/sec, ~0.0190 J/pos (~1.17× throughput, ~10% lower board energy). Batch 32: LuxiEdge ~43,900 pos/sec, ~0.0158 J/pos vs vLLM ~37,100 pos/sec, ~0.0182 J/pos (~1.18× throughput, ~13% lower board energy). Prompt positions per second are not full chat decode tokens per second. Board energy via NVML - not facility, wall-plug, or PUE-adjusted energy. This is not a multi-tenant serving-stack or every-workload claim. This measurement is not a TESTfort or third-party evaluation. Public brief: H2H Prefill Energy Brief.
**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.
Full methods and downloads: Proof.
Measurement scope: one NVIDIA H100 80GB-class GPU · Qwen2-7B-Instruct FP16 · batches 16 and 32 · sequence length 128 · matched prefill positions · one generated token per iteration · NVML GPU-board energy. Prompt positions per second are not full chat decode tokens per second. Not facility, wall-plug, or PUE-adjusted energy. Not a multi-tenant serving-stack or every-workload result. Detailed methods →
Faithful Llama 3.1 resident inference
Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.
Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.
A three-legged stool
Pick the path that fits you. Every path can end in shareable proof.
Quant & research
Determinism, audit trails, free-ride paths, long-context memory - for people who need methods, not slogans.
AI & data centers
Power caps, density, measured joules-per-token under load, and a commercial path to evaluation.
We need your help
You don’t need a PhD. Share the energy story. Ask AI providers and data centers why they aren’t cutting waste.