Deterministic execution · measured energy · public evidence
More useful compute from every watt.
LuxiEdge is building technology that helps computers do more useful work with less electricity. Our systems are designed to produce consistent, verifiable results while delivering the speed real-world applications require. We measure energy use and performance against established systems, so every improvement can be shown, not merely claimed. Today, that work includes LuxiQuant, our deterministic engine for numerical and quantitative workloads, and our measured H100 prefill technology for more energy-efficient AI inference. LuxiEdge is patent pending.
~1.17-1.18× throughput · ~10-14% lower board J/position vs vLLM†
Batch 16 and 32 · Qwen2-7B-Instruct FP16 · H100 · Version 100
Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.
†Version 100 internal measurement, July 2026. Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, one generated token per iteration. Batch 16: ~1.17× throughput, ~10% lower board energy vs vLLM. Batch 32: ~1.18× throughput, ~13% lower board energy vs vLLM. Exact per-batch values are being reconciled; public pages present ratio ranges only. Prompt positions per second are not full chat decode tokens per second. Board energy via NVML - not facility, wall-plug, or PUE-adjusted energy. This is not a multi-tenant serving-stack or every-workload claim. This measurement is not a TESTfort or third-party evaluation. Public brief: H2H Prefill Energy Brief.
**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.
Full methods and downloads: Proof.
Measurement scope: one NVIDIA H100 80GB-class GPU · Qwen2-7B-Instruct FP16 · batches 16 and 32 · sequence length 128 · matched prefill positions · one generated token per iteration · NVML GPU-board energy. Prompt positions per second are not full chat decode tokens per second. Not facility, wall-plug, or PUE-adjusted energy. Not a multi-tenant serving-stack or every-workload result. Detailed methods →
Two offers, both measured against your workload
Public evidence and published methods stay free. Hands-on engagement is paid, scoped, and reported against your own workload, hardware, and reference runtime.
LuxiQuant
A bit-exact deterministic numerical engine - not LLM weight quantization. Evaluated by TestFort on a seven-function workload with identical output hashes across tested platforms. Design-partner pilots measure it against your numerical workload.
H100 Prefill Evaluation
A scoped, paid head-to-head measurement of prompt-position throughput and board energy per position against your reference runtime, under an agreed measurement contract on H100 hardware.
What measured efficiency is worth to a buyer
Energy and speed are not abstractions. Our measurement program quantifies, per workload:
Fewer joules per useful unit
Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.
More work per GPU-hour
Higher throughput on the same board means more requests served per hour of capacity you already pay for.
More capacity under a power cap
When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.
Faster completion, deferred capital
Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.
Internal Research
The work below is internal research with the maturity labels shown. It is not a product offer, and no production-serving claims are made for it.
Faithful Llama 3.1 resident inference
Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.
Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.
Public action - not a sales pitch
We need your help
AI infrastructure already consumes a large and fast-growing share of electricity, and the demand curve is still rising. Before new power plants, new transmission lines, and higher community utility bills are treated as inevitable, providers and data centers should measure energy per unit of useful work and cut computational waste. You do not need to be an engineer to push that measurement into the open.
Share the evidence
Our methods and numbers are public. One share is a vote for measured efficiency over slogans.
Ask for joules per useful unit
Ask your AI provider or data center whether they publish measured energy per unit of useful work - hardware, workload, and power boundary included.
Contact decision makers
Efficiency is a public interest. Utilities, regulators, and local officials should ask for measurement before approving new power demand.
Connect and support testing
Introduce LuxiEdge to operators and independent evaluators, and support third-party testing of energy-per-work claims - ours included.