Three audiences · one mission

AI that uses less electricity - and proves it.

9.15% less board energy · 91.78% of vLLM batch-invariant throughput†

3.10% less board energy · 80.60% of default vLLM throughput†

Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.

Third-party tested on H100. Lower board energy per prefill than vLLM. Faithful Qwen2-7B CUDA correctness gates passed.

9.15%
Less board energy vs batch-invariant vLLM†
91.78% of batch-invariant throughput · V99
3.10%
Less board energy vs default vLLM†
80.60% of default throughput · V99
Full 28-layer
Faithful Qwen2-7B CUDA path on H100
1.0
HF/CUDA greedy-token match rate**

†TESTfort third-party measurement (Version 99, 2026-07-23). Qwen2-7B-Instruct packed prefill, batch 16, sequence length 128, one NVIDIA H100 80GB. Work unit: prefill positions - not decode tokens or full serving throughput. LuxiEdge medians: 28,374.7 positions/sec, 0.018718 board J/position. Default vLLM medians: 35,203.1 positions/sec, 0.019316 board J/position. LuxiEdge achieved 80.60% of default vLLM throughput and used 3.10% less board energy per position. This is not a throughput win. Formal signed narrative pending; raw artifacts and technician attestation available. Board energy via NVML - not facility or wall-plug energy. Results apply to the tested configuration and workload.

**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.

Full methods and downloads: Proof.

Measurement scope: one NVIDIA H100 80GB · Qwen2-7B-Instruct · batch 16 · sequence 128 packed prefill positions · NVML GPU-board energy. Not decode tokens, full serving throughput, facility power, or wall-plug energy. Results apply to the tested configuration. Detailed methods →

Internal H100 milestone - independent validation pending

Faithful Llama 3.1 resident inference

Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.

Resident
Weights and KV state
GPU-resident path
BF16
Tensor Core path
H100 NVL
Exact
Batch-to-serial agreement on tested gates
Primary p32 and p128 gates
Internal
Performance and energy validation pending

Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.

A three-legged stool

Pick the path that fits you. Every path can end in shareable proof.

Quant & research

Determinism, audit trails, free-ride paths, long-context memory - for people who need methods, not slogans.

Science path →

AI & data centers

Power caps, density, measured joules-per-token under load, and a commercial path to evaluation.

Operator path →

We need your help

You don’t need a PhD. Share the energy story. Ask AI providers and data centers why they aren’t cutting waste.

Public path →

One layer of a larger architecture

LuxiEdge is the first public product in the Luxi energy architecture for AI compute. The full system spans deterministic computation, GPU execution, request scheduling, load shaping, and facility power control - each layer targeting a different source of wasted energy, each with its own measured evidence and maturity label.