Three audiences · one mission

Energy-efficient AI inference with public H100 evidence.

~1.17-1.18× throughput · ~10-14% lower board J/position vs vLLM†

Batch 16 and 32 · Qwen2-7B-Instruct FP16 · H100 · Version 100

Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.

Higher throughput and lower board energy than the tested vLLM stack on this prefill workload. Faithful Qwen2-7B CUDA correctness gates passed.

~1.17-1.18×
Higher throughput vs vLLM†
Batch 16: ~41,700 vs ~35,800 pos/sec · V100
~10-14%
Lower board J/position vs vLLM†
Batch 16: ~0.0171 vs ~0.0190 J/pos · V100
1.0
Dual-run determinism score†
1.0
HF/CUDA greedy-token match rate**

†Version 100 internal measurement, July 2026. Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, one generated token per iteration. Batch 16: LuxiEdge ~41,700 pos/sec, ~0.0171 J/pos vs vLLM ~35,800 pos/sec, ~0.0190 J/pos (~1.17× throughput, ~10% lower board energy). Batch 32: LuxiEdge ~43,900 pos/sec, ~0.0158 J/pos vs vLLM ~37,100 pos/sec, ~0.0182 J/pos (~1.18× throughput, ~13% lower board energy). Prompt positions per second are not full chat decode tokens per second. Board energy via NVML - not facility, wall-plug, or PUE-adjusted energy. This is not a multi-tenant serving-stack or every-workload claim. This measurement is not a TESTfort or third-party evaluation. Public brief: H2H Prefill Energy Brief.

**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.

Full methods and downloads: Proof.

Measurement scope: one NVIDIA H100 80GB-class GPU · Qwen2-7B-Instruct FP16 · batches 16 and 32 · sequence length 128 · matched prefill positions · one generated token per iteration · NVML GPU-board energy. Prompt positions per second are not full chat decode tokens per second. Not facility, wall-plug, or PUE-adjusted energy. Not a multi-tenant serving-stack or every-workload result. Detailed methods →

Internal H100 milestone - independent validation pending

Faithful Llama 3.1 resident inference

Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.

Resident
Weights and KV state
GPU-resident path
BF16
Tensor Core path
H100 NVL
Exact
Batch-to-serial agreement on tested gates
Primary p32 and p128 gates
Internal
Performance and energy validation pending

Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.

A three-legged stool

Pick the path that fits you. Every path can end in shareable proof.

Quant & research

Determinism, audit trails, free-ride paths, long-context memory - for people who need methods, not slogans.

Science path →

AI & data centers

Power caps, density, measured joules-per-token under load, and a commercial path to evaluation.

Operator path →

We need your help

You don’t need a PhD. Share the energy story. Ask AI providers and data centers why they aren’t cutting waste.

Public path →

One layer of a larger architecture

LuxiEdge is the first public product in the Luxi energy architecture for AI compute. The full system spans deterministic computation, GPU execution, request scheduling, load shaping, and facility power control - each layer targeting a different source of wasted energy, each with its own measured evidence and maturity label.