Deterministic execution · measured energy · public evidence

More useful compute from every watt.

LuxiEdge is building technology that helps computers do more useful work with less electricity. Our systems are designed to produce consistent, verifiable results while delivering the speed real-world applications require. We measure energy use and performance against established systems, so every improvement can be shown, not merely claimed. Today, that work includes LuxiQuant, our deterministic engine for numerical and quantitative workloads, and our measured H100 prefill technology for more energy-efficient AI inference. LuxiEdge is patent pending.

What LuxiEdge offers today

Two separate lanes of work, each described in plain language and each backed by the evidence on our Proof page.

LuxiQuant: a deterministic numerical engine

LuxiQuant computes numerical and quantitative results that repeat exactly, run after run. It was independently evaluated by TestFort QA Lab, which reported identical output hashes across the tested runs and platforms. It is separate from transformer AI inference. Paid design-partner pilot work is open today, measured against your own numerical workload.

LuxiQuant pilots →

H100 inference: a bounded evaluation lane

Our faithful full-model CUDA path runs the complete model and passes its defined reference-agreement tests against established implementations. We can measure prefill energy and speed head-to-head against your reference workload on H100 hardware under an agreed measurement plan. True-autoregressive decode remains internal research while its current speed and energy are remeasured.

The evaluation offer →

Where the AI inference work stands

Here is the honest current picture, in order.

An independent test lab measured our H100 prefill technology against default vLLM, a widely used serving system. On that workload, our system used 3.10% less GPU-board energy per prefill position, while processing 80.60% as many positions per second as vLLM. In other words, we used measurably less energy per unit of work, and we were slower on that test.

Later internal prefill tests looked better than that independent result. But our exact internal results conflicted between runs, so we withdrew the headline numbers and are running a frozen matched rerun before publishing a replacement.

Separately, a numerical audit of our true-autoregressive GPU path found that required mathematics had been omitted or substituted. We corrected the path, and it now passes the defined reference-agreement audit. Its current decode speed and energy have not yet been remeasured.

Put simply: the independently measured energy signal is real, the corrected inference path is working under defined correctness tests, and the current performance result is still being established.

Full methods, numbers, and limitations: Proof.

Two offers, both measured against your workload

Public evidence and published methods stay free. Hands-on engagement is paid, scoped, and reported against your own workload, hardware, and reference runtime.

Paid Pilot Open

LuxiQuant

A bit-exact deterministic numerical engine - not LLM weight quantization. Evaluated by TestFort on a seven-function workload with identical output hashes across tested platforms. Design-partner pilots measure it against your numerical workload.

LuxiQuant pilots →

Scoped Evaluation Open

H100 Prefill Evaluation

A scoped, paid head-to-head measurement of prompt-position throughput and board energy per position against your reference runtime, under an agreed measurement contract on H100 hardware.

The evaluation offer →

What measured efficiency is worth to a buyer

Energy and speed are not abstractions. Our measurement program quantifies, per workload:

Fewer joules per useful unit

Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.

More work per GPU-hour

Higher throughput on the same board means more requests served per hour of capacity you already pay for.

More capacity under a power cap

When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.

Faster completion, deferred capital

Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.

These are quantities we measure, not savings we promise. Every engagement begins with a defined workload, a reference runtime, and an agreed measurement contract. We report what we measure - including when we do not win.

Internal Research

The work below is internal research with the maturity labels shown. It is not a product offer, and no production-serving claims are made for it.

Internal H100 milestone - independent validation pending

Faithful Llama 3.1 resident inference

Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.

Resident
Weights and KV state
GPU-resident path
BF16
Tensor Core path
H100 NVL
Exact
Batch-to-serial agreement on tested gates
Primary p32 and p128 gates
Internal
Performance and energy validation pending

Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.

Research status: true autoregressive decode exists internally at LuxiEdge. It is not offered as a production serving product, and no decode performance numbers are published on this page. Matched-vLLM comparisons against production serving engines are pending.

One layer of a larger architecture

LuxiEdge is the first public product in the Luxi energy architecture for AI compute. The full system spans deterministic computation, GPU execution, request scheduling, load shaping, and facility power control - each layer targeting a different source of wasted energy, each with its own measured evidence and maturity label.

Public action - not a sales pitch

We need your help

AI systems and data centers create enormous value, and their electricity use is growing quickly. That makes computational efficiency a legitimate public issue. LuxiEdge is not asking people to oppose data centers or technological progress. We are asking operators, researchers, customers, utilities, public officials, and communities to measure energy per unit of useful work, share evidence, and help better methods spread. Everyone can take part in that conversation.

Learn and share the evidence

Our methods and numbers are public, corrections included. Sharing them helps measured efficiency spread faster than slogans.

Ask how energy is measured

Ask your AI provider or data center how they measure energy per unit of useful work - hardware, workload, and power boundary included. It is a fair question, and good operators welcome it.

Talk with decision makers

Efficiency is a public interest. You can encourage utilities, regulators, and local officials to include software efficiency and energy-per-work measurement in their planning conversations.

Connect and support testing

Introduce LuxiEdge to operators, researchers, and independent evaluators, and support third-party testing of energy-per-work claims - ours included.