Two engines · deterministic compute you can prove · energy you can measure

Deterministic compute you can prove. Energy you can measure.

Luxi has two engines. LuxiQuant is a deterministic numerical and quantitative engine, ready for pilot today. The Luxi Inference Engine runs transformer workloads in two stages: prefill, which reads the prompt, and decode/generation, which produces output tokens. Prefill is available today as a controlled evaluation; decode/generation matches the reference next-token result under the defined numerical audit and is on the road to serving. LuxiEdge is patent pending.

The two engines

Two separate engines, each described in plain language and each backed by the evidence on our Evidence page. Energy is the commercial value we measure; determinism is the floor everything stands on.

Ready for pilot

LuxiQuant - deterministic numerical engine

LuxiQuant computes numerical and quantitative results that repeat exactly, run after run, with a SHA-256 receipt attached to every output. It was independently evaluated by TestFort QA Lab, which reported identical output hashes across the tested runs and platforms. It serves quant finance, risk, and any pipeline where results must be reproduced and proven. Paid design-partner pilots are open today, measured against your own numerical workload.

Deterministic numerical computing for servers and constrained ARM systems.

Pilot LuxiQuant →

Prefill: evaluation open · Decode: in development

Luxi Inference Engine - transformer workloads in stages

Prefill reads the prompt. It has an independent measured baseline and a stronger internal matched result, each clearly labeled, and is available as a controlled, paid evaluation on your workload. Decode/generation produces output tokens and matches the reference next-token result under the defined numerical audit; serving readiness is the next milestone.

Request an Inference Prefill Evaluation →

Where the Inference Engine stands, stage by stage

Here is the honest current picture, by stage and evidence class.

Internal absolute prefill

Approximately 44,860 prefill positions per second at approximately 0.01532 GPU-board joules per prefill position. Scope: one NVIDIA H100 80GB HBM3; Qwen2-7B-Instruct-class weights; sequence length 128; batch 72; dual-GEMM; Flash attention; device-resident FP16 path; median of five 15-second runs. Internal, absolute - not a matched vLLM comparison.

Absolute evidence pack →

Internal matched comparison

At batch 16, approximately 1.18 times the tested vLLM throughput at approximately 12 percent lower GPU-board joules per prefill position. Internal, matched, same-GPU comparison - this ratio belongs to the batch-16 comparison, not to the batch-72 absolute result.

Matched evidence pack →

Absolute batch ≠ matched batch. The independent TESTfort baseline is separate. Board energy is GPU-board energy measured through NVML, not wall-plug or facility energy. Prefill positions are not generated tokens.

Prefill - independent baseline

Independently measured by a third-party test lab: 3.10% lower GPU-board energy per prefill position than default vLLM at 80.60% of its throughput on that workload - a verified energy edge, with the complete throughput tradeoff published alongside it.

Decode/generation

Decode/generation produces output tokens and now matches the reference next-token result under the defined numerical audit. The next serving-readiness milestone is locking end-to-end quality, throughput, and board energy, after which decode becomes a serving offer.

What that means for buyers

You can pilot LuxiQuant today, and you can commission a controlled prefill evaluation on your own workload today. Decode is progressing toward a serving offer, with quality, throughput, and board energy being locked next.

Put simply: the independently measured prefill energy signal is real, the internal matched and absolute results are stronger still - each reported separately - and decode matches the reference next-token result under the defined numerical audit on its way to a serving offer.

Full methods, numbers, and limitations, organized by engine and stage: Evidence.

What measured efficiency is worth to a buyer

Energy and speed are not abstractions. Our measurement program quantifies, per workload:

Fewer joules per useful unit

Every processed position or numerical result carries an energy cost. We measure it head-to-head against your reference runtime, on your hardware.

More work per GPU-hour

Higher throughput on the same board means more requests served per hour of capacity you already pay for.

More capacity under a power cap

When the facility power cap is fixed, lower joules per unit of useful work is the only way to grow output without new power.

Faster completion, deferred capital

Faster prefill shortens every request that depends on it. Sustained efficiency gains can reduce GPU-count growth and delay new hardware and power purchases.

These are quantities we measure, not savings we promise. Every engagement begins with a defined workload, a reference runtime, and an agreed measurement contract. You receive the complete measured tradeoff, and you can evaluate LuxiEdge on your own workload before deciding anything.

One layer of a larger architecture

The two engines are the first public products in the Luxi energy architecture for AI compute. The full system spans deterministic computation, GPU execution, request scheduling, load shaping, and facility power control - each layer targeting a different source of wasted energy, each with its own measured evidence and maturity label.

Public action - not a sales pitch

We need your help

AI systems and data centers create enormous value, and their electricity use is growing quickly. That makes computational efficiency a legitimate public issue. LuxiEdge is not asking people to oppose data centers or technological progress. We are asking operators, researchers, customers, utilities, public officials, and communities to measure energy per unit of useful work, share evidence, and help better methods spread. Everyone can take part in that conversation.

Learn and share the evidence

Our methods and numbers are public and versioned - every figure traces to its protocol. Sharing them helps measured efficiency spread faster than slogans.

Ask how energy is measured

Ask your AI provider or data center how they measure energy per unit of useful work - hardware, workload, and power boundary included. It is a fair question, and good operators welcome it.

Talk with decision makers

Efficiency is a public interest. You can encourage utilities, regulators, and local officials to include software efficiency and energy-per-work measurement in their planning conversations.

Connect and support testing

Introduce LuxiEdge to operators, researchers, and independent evaluators, and support third-party testing of energy-per-work claims - ours included.