Two engines · deterministic compute you can prove · energy you can measure
Deterministic compute you can prove. Energy you can measure.
Luxi has two engines. LuxiQuant is a deterministic numerical and quantitative engine, ready for pilot today. The Luxi Inference Engine runs transformer workloads in two stages: prefill, which reads the prompt, and decode/generation, which produces output tokens. Prefill is available today as a controlled evaluation; decode/generation matches the reference next-token result under the defined numerical audit and is on the road to serving. LuxiEdge is patent pending.
The two engines
Two separate engines, each described in plain language and each backed by the evidence on our Evidence page. Energy is the commercial value we measure; determinism is the floor everything stands on.
LuxiQuant - deterministic numerical engine
LuxiQuant computes numerical and quantitative results that repeat exactly, run after run, with a SHA-256 receipt attached to every output. It was independently evaluated by TestFort QA Lab, which reported identical output hashes across the tested runs and platforms. It serves quant finance, risk, and any pipeline where results must be reproduced and proven. Paid design-partner pilots are open today, measured against your own numerical workload.
Deterministic numerical computing for servers and constrained ARM systems.
Luxi Inference Engine - transformer workloads in stages
Prefill reads the prompt. It has an independent measured baseline and a stronger internal matched result, each clearly labeled, and is available as a controlled, paid evaluation on your workload. Decode/generation produces output tokens and matches the reference next-token result under the defined numerical audit; serving readiness is the next milestone.
Where the Inference Engine stands, stage by stage
Here is the honest current picture, by stage and evidence class.
Internal absolute prefill
Approximately 44,860 prefill positions per second at approximately 0.01532 GPU-board joules per prefill position. Scope: one NVIDIA H100 80GB HBM3; Qwen2-7B-Instruct-class weights; sequence length 128; batch 72; dual-GEMM; Flash attention; device-resident FP16 path; median of five 15-second runs. Internal, absolute - not a matched vLLM comparison.
Internal matched comparison
At batch 16, approximately 1.18 times the tested vLLM throughput at approximately 12 percent lower GPU-board joules per prefill position. Internal, matched, same-GPU comparison - this ratio belongs to the batch-16 comparison, not to the batch-72 absolute result.
Prefill - independent baseline
Independently measured by a third-party test lab: 3.10% lower GPU-board energy per prefill position than default vLLM at 80.60% of its throughput on that workload - a verified energy edge, with the complete throughput tradeoff published alongside it.
Decode/generation
Decode/generation produces output tokens and now matches the reference next-token result under the defined numerical audit. The next serving-readiness milestone is locking end-to-end quality, throughput, and board energy, after which decode becomes a serving offer.
What that means for buyers
You can pilot LuxiQuant today, and you can commission a controlled prefill evaluation on your own workload today. Decode is progressing toward a serving offer, with quality, throughput, and board energy being locked next.
Full methods, numbers, and limitations, organized by engine and stage: Evidence.
What measured efficiency is worth to a buyer
Energy and speed are not abstractions. Our measurement program quantifies, per workload:
Fewer joules per useful unit
Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.
More work per GPU-hour
Higher throughput on the same board means more requests served per hour of capacity you already pay for.
More capacity under a power cap
When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.
Faster completion, deferred capital
Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.
Public action - not a sales pitch
We need your help
AI systems and data centers create enormous value, and their electricity use is growing quickly. That makes computational efficiency a legitimate public issue. LuxiEdge is not asking people to oppose data centers or technological progress. We are asking operators, researchers, customers, utilities, public officials, and communities to measure energy per unit of useful work, share evidence, and help better methods spread. Everyone can take part in that conversation.
Learn and share the evidence
Our methods and numbers are public and versioned - every figure traces to its protocol. Sharing them helps measured efficiency spread faster than slogans.
Ask how energy is measured
Ask your AI provider or data center how they measure energy per unit of useful work - hardware, workload, and power boundary included. It is a fair question, and good operators welcome it.
Talk with decision makers
Efficiency is a public interest. You can encourage utilities, regulators, and local officials to include software efficiency and energy-per-work measurement in their planning conversations.
Connect and support testing
Introduce LuxiEdge to operators, researchers, and independent evaluators, and support third-party testing of energy-per-work claims - ours included.