Deterministic execution · measured energy · public evidence
More useful compute from every watt.
LuxiEdge is building technology that helps computers do more useful work with less electricity. Our systems are designed to produce consistent, verifiable results while delivering the speed real-world applications require. We measure energy use and performance against established systems, so every improvement can be shown, not merely claimed. Today, that work includes LuxiQuant, our deterministic engine for numerical and quantitative workloads, and our measured H100 prefill technology for more energy-efficient AI inference. LuxiEdge is patent pending.
What LuxiEdge offers today
Two separate lanes of work, each described in plain language and each backed by the evidence on our Proof page.
LuxiQuant: a deterministic numerical engine
LuxiQuant computes numerical and quantitative results that repeat exactly, run after run. It was independently evaluated by TestFort QA Lab, which reported identical output hashes across the tested runs and platforms. It is separate from transformer AI inference. Paid design-partner pilot work is open today, measured against your own numerical workload.
H100 inference: a bounded evaluation lane
Our faithful full-model CUDA path runs the complete model and passes its defined reference-agreement tests against established implementations. We can measure prefill energy and speed head-to-head against your reference workload on H100 hardware under an agreed measurement plan. True-autoregressive decode remains internal research while its current speed and energy are remeasured.
Where the AI inference work stands
Here is the honest current picture, in order.
An independent test lab measured our H100 prefill technology against default vLLM, a widely used serving system. On that workload, our system used 3.10% less GPU-board energy per prefill position, while processing 80.60% as many positions per second as vLLM. In other words, we used measurably less energy per unit of work, and we were slower on that test.
Later internal prefill tests looked better than that independent result. But our exact internal results conflicted between runs, so we withdrew the headline numbers and are running a frozen matched rerun before publishing a replacement.
Separately, a numerical audit of our true-autoregressive GPU path found that required mathematics had been omitted or substituted. We corrected the path, and it now passes the defined reference-agreement audit. Its current decode speed and energy have not yet been remeasured.
Full methods, numbers, and limitations: Proof.
Two offers, both measured against your workload
Public evidence and published methods stay free. Hands-on engagement is paid, scoped, and reported against your own workload, hardware, and reference runtime.
LuxiQuant
A bit-exact deterministic numerical engine - not LLM weight quantization. Evaluated by TestFort on a seven-function workload with identical output hashes across tested platforms. Design-partner pilots measure it against your numerical workload.
H100 Prefill Evaluation
A scoped, paid head-to-head measurement of prompt-position throughput and board energy per position against your reference runtime, under an agreed measurement contract on H100 hardware.
What measured efficiency is worth to a buyer
Energy and speed are not abstractions. Our measurement program quantifies, per workload:
Fewer joules per useful unit
Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.
More work per GPU-hour
Higher throughput on the same board means more requests served per hour of capacity you already pay for.
More capacity under a power cap
When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.
Faster completion, deferred capital
Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.
Internal Research
The work below is internal research with the maturity labels shown. It is not a product offer, and no production-serving claims are made for it.
Faithful Llama 3.1 resident inference
Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.
Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.
Public action - not a sales pitch
We need your help
AI systems and data centers create enormous value, and their electricity use is growing quickly. That makes computational efficiency a legitimate public issue. LuxiEdge is not asking people to oppose data centers or technological progress. We are asking operators, researchers, customers, utilities, public officials, and communities to measure energy per unit of useful work, share evidence, and help better methods spread. Everyone can take part in that conversation.
Learn and share the evidence
Our methods and numbers are public, corrections included. Sharing them helps measured efficiency spread faster than slogans.
Ask how energy is measured
Ask your AI provider or data center how they measure energy per unit of useful work - hardware, workload, and power boundary included. It is a fair question, and good operators welcome it.
Talk with decision makers
Efficiency is a public interest. You can encourage utilities, regulators, and local officials to include software efficiency and energy-per-work measurement in their planning conversations.
Connect and support testing
Introduce LuxiEdge to operators, researchers, and independent evaluators, and support third-party testing of energy-per-work claims - ours included.