Deterministic execution · measured energy · public evidence

More useful compute from every watt.

LuxiEdge is building technology that helps computers do more useful work with less electricity. Our systems are designed to produce consistent, verifiable results while delivering the speed real-world applications require. We measure energy use and performance against established systems, so every improvement can be shown, not merely claimed. Today, that work includes LuxiQuant, our deterministic engine for numerical and quantitative workloads, and our measured H100 prefill technology for more energy-efficient AI inference. LuxiEdge is patent pending.

~1.17-1.18× throughput · ~10-14% lower board J/position vs vLLM†

Batch 16 and 32 · Qwen2-7B-Instruct FP16 · H100 · Version 100

Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.

~1.17-1.18×
Higher throughput vs vLLM†
Ratio range across batches 16 and 32 · V100
~10-14%
Lower board J/position vs vLLM†
Ratio range across batches 16 and 32 · V100
Pass / fail
Determinism contract - we declare exactly what repeats
Numerical result, token, tensor, logit, KV state, or receipt
1.0
HF/CUDA greedy-token match rate**

†Version 100 internal measurement, July 2026. Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, one generated token per iteration. Batch 16: ~1.17× throughput, ~10% lower board energy vs vLLM. Batch 32: ~1.18× throughput, ~13% lower board energy vs vLLM. Exact per-batch values are being reconciled; public pages present ratio ranges only. Prompt positions per second are not full chat decode tokens per second. Board energy via NVML - not facility, wall-plug, or PUE-adjusted energy. This is not a multi-tenant serving-stack or every-workload claim. This measurement is not a TESTfort or third-party evaluation. Public brief: H2H Prefill Energy Brief.

**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.

Full methods and downloads: Proof.

Measurement scope: one NVIDIA H100 80GB-class GPU · Qwen2-7B-Instruct FP16 · batches 16 and 32 · sequence length 128 · matched prefill positions · one generated token per iteration · NVML GPU-board energy. Prompt positions per second are not full chat decode tokens per second. Not facility, wall-plug, or PUE-adjusted energy. Not a multi-tenant serving-stack or every-workload result. Detailed methods →

Two offers, both measured against your workload

Public evidence and published methods stay free. Hands-on engagement is paid, scoped, and reported against your own workload, hardware, and reference runtime.

Paid Pilot Open

LuxiQuant

A bit-exact deterministic numerical engine - not LLM weight quantization. Evaluated by TestFort on a seven-function workload with identical output hashes across tested platforms. Design-partner pilots measure it against your numerical workload.

LuxiQuant pilots →

Scoped Evaluation Open

H100 Prefill Evaluation

A scoped, paid head-to-head measurement of prompt-position throughput and board energy per position against your reference runtime, under an agreed measurement contract on H100 hardware.

The evaluation offer →

What measured efficiency is worth to a buyer

Energy and speed are not abstractions. Our measurement program quantifies, per workload:

Fewer joules per useful unit

Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.

More work per GPU-hour

Higher throughput on the same board means more requests served per hour of capacity you already pay for.

More capacity under a power cap

When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.

Faster completion, deferred capital

Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.

These are quantities we measure, not savings we promise. Every engagement begins with a defined workload, a reference runtime, and an agreed measurement contract. We report what we measure - including when we do not win.

Internal Research

The work below is internal research with the maturity labels shown. It is not a product offer, and no production-serving claims are made for it.

Internal H100 milestone - independent validation pending

Faithful Llama 3.1 resident inference

Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.

Resident
Weights and KV state
GPU-resident path
BF16
Tensor Core path
H100 NVL
Exact
Batch-to-serial agreement on tested gates
Primary p32 and p128 gates
Internal
Performance and energy validation pending

Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.

Research status: true autoregressive decode exists internally at LuxiEdge. It is not offered as a production serving product, and no decode performance numbers are published on this page. Matched-vLLM comparisons against production serving engines are pending.

One layer of a larger architecture

LuxiEdge is the first public product in the Luxi energy architecture for AI compute. The full system spans deterministic computation, GPU execution, request scheduling, load shaping, and facility power control - each layer targeting a different source of wasted energy, each with its own measured evidence and maturity label.

Public action - not a sales pitch

We need your help

AI infrastructure already consumes a large and fast-growing share of electricity, and the demand curve is still rising. Before new power plants, new transmission lines, and higher community utility bills are treated as inevitable, providers and data centers should measure energy per unit of useful work and cut computational waste. You do not need to be an engineer to push that measurement into the open.

Share the evidence

Our methods and numbers are public. One share is a vote for measured efficiency over slogans.

Ask for joules per useful unit

Ask your AI provider or data center whether they publish measured energy per unit of useful work - hardware, workload, and power boundary included.

Contact decision makers

Efficiency is a public interest. Utilities, regulators, and local officials should ask for measurement before approving new power demand.

Connect and support testing

Introduce LuxiEdge to operators and independent evaluators, and support third-party testing of energy-per-work claims - ours included.