By engine · by stage · by evidence class

Every number, its engine, its stage, and its limits.

Luxi has two engines: LuxiQuant, a deterministic numerical engine ready for pilot, and the Luxi Inference Engine, which runs transformer workloads in stages - prefill and decode/generation. Every material claim on this site is registered here under its engine and stage with its work unit, comparator, hardware and model scope, energy boundary, determinism or correctness contract, maturity label, limitations, and public data-pack link where one exists.

How to read our claims

Claims are organized by engine (LuxiQuant or the Luxi Inference Engine), by stage for inference (prefill or decode/generation), and into three evidence classes that never blend. Nothing here is ever merged into one system-wide win. Historical tags such as Version 99, 100, 101, or 102 are internal milestone tags, not product names; where they matter for provenance, they appear only in footnotes.

Independent Evidence

Measured or operated by a third party (TestFort QA Lab) under its own protocol, with dated reports. The strongest evidence class on this page.

Internal Matched Evidence

Internal head-to-head measurement against a named comparator under a stated protocol, or correctness acceptance against a reference implementation. Independent validation is noted wherever it exists; everything else is labeled internal.

Research Evidence

Internal research-path measurements. Directional evidence about architecture and energy behavior - not product claims, not production-serving evidence.

Maturity labels

Independent Evidence and Internal Research mark proof status. Paid Pilot Open (LuxiQuant) and Scoped Evaluation Open (H100 prefill) mark commercial engagements that measure your workload under contract.

How to read our numbers. Independent evidence is measured by a third party; internal evidence is measured by us against a named comparator - the two are always reported separately. Prefill results count evaluated prompt positions per second; decode results count generated tokens - the two units are never interchanged. Energy is GPU-board energy via NVML, measured on the stated workload and configuration; facility, wall-plug, and PUE-adjusted energy are different quantities we do not claim. Every result applies to its stated workload and configuration.

Current situation

Where the evidence stands right now

LuxiQuant, the deterministic numerical engine, was independently evaluated by TestFort QA Lab and is available for paid pilots today.

Inference Engine - prefill: three separate evidence rungs, in order. The independent TESTfort baseline (July 23, 2026) verified lower GPU-board energy per prefill position than the tested vLLM stack. The internal matched comparison measured roughly 1.18x the tested vLLM throughput at roughly 12% lower board joules per position at batch 16. The internal absolute prefill champion (August 7, 2026) measured roughly 44,860 positions per second at roughly 0.01532 board joules per position at batch 72, with no vLLM arm attached. Each rung is reported separately below. Prefill is available today as a controlled evaluation.

Inference Engine - decode/generation: decode produces output tokens and matches the reference next-token result under the defined numerical audit. The next serving-readiness milestone is locking end-to-end quality, throughput, and board energy.

LuxiQuant · Independent Evidence

LuxiQuant: independently evaluated deterministic numerical engine

TestFort QA Lab independently evaluated LuxiQuant, LuxiEdge's deterministic numerical engine, on NVIDIA H100 hardware. This third-party numerical validation belongs to LuxiQuant and is never applied to the Inference Engine's transformer results.

286.94B
Aggregate operations per second
Tested seven-function workload
331.13B
Peak operations per second
Tested square-root operation
444.4T
Reported operations
One-hour load test
1.47 ms
p95 API latency
Reported 200-user load configuration

Reliability

  • No failed requests in the specified load test
  • Average GPU power of approximately 117.2 W during the reported load test

Determinism

  • Five GPU and five CPU runs produced identical hashes for the tested workload
Scope: TestFort independently evaluated a defined deterministic numeric workload on an NVIDIA H100 SXM. These results apply to the tested workload and configuration and are separate from LuxiEdge's transformer research benchmarks.

TestFort evaluation report - December 2025 (PDF) →
TestFort packed-prefill evaluation report - July 2026 (PDF) →

The December 2025 report covers the deterministic numeric engine evaluation (seven-function workload, hash consistency, GPU endurance). The July 2026 report covers independent evaluation of LuxiEdge deterministic packed prefill execution versus vLLM on NVIDIA H100 (Version 99, 2026-07-23 results). These are two separate evaluations by TestFort QA Lab.

Luxi Inference Engine · Prefill · Independent Evidence

Independent prefill baseline - July 23, 2026

TESTfort third-party measurement completed 2026-07-23, with the validation report, raw artifacts, and technician attestation all available. Workload: Qwen2-7B-Instruct packed prefill, batch 16, sequence length 128, one NVIDIA H100 80GB, measured in prefill positions. Preserved as dated independent evidence, reported separately from the later internal prefill measurement below.

3.10%
Lower board J/position vs default vLLM†
0.018718 vs 0.019316 J/position
9.15%
Lower board J/position vs batch-invariant vLLM†
0.018718 vs 0.020604 J/position
80.60%
Of default vLLM throughput†
28,374.7 vs 35,203.1 positions/sec
91.78%
Of batch-invariant vLLM throughput†
28,374.7 vs 30,914.3 positions/sec
ConfigurationThroughput (positions/sec)Board J/positionThroughput ratioEnergy ratio
LuxiEdge prefill build‡28,374.70.018718--
vLLM (default)35,203.10.019316LuxiEdge 80.60% of vLLMLuxiEdge 3.10% lower
vLLM (batch-invariant)30,914.30.020604LuxiEdge 91.78% of vLLMLuxiEdge 9.15% lower
Independently verified: 3.10% lower board joules per prefill position than default vLLM (9.15% lower than the batch-invariant configuration), at 80.6% of default vLLM throughput on this workload. The complete measured tradeoff - energy and throughput - is published in full above.
Scope: packed prefill, Qwen2-7B-Instruct, batch 16, sequence length 128, one H100 80GB; board energy via NVML (see the reading guide above). The tested LuxiEdge path had stable pinned-environment fingerprints, zero measured cross-sequence contamination, and zero FlashAttention fallback. Determinism findings apply to the tested hardware and backend configuration.

†Medians from the TESTfort measurement of 2026-07-23. Prefill positions only - not a full serving or decode comparison. Board energy is not facility or wall-plug electricity. ‡The tested build carries the internal milestone tag Version 99; that tag is an internal milestone, not a product name.

Report: TestFort validation report - July 2026 (PDF) → · Evidence pack →

Luxi Inference Engine · Prefill · Internal Matched Evidence

Internal matched comparison - batch 16

Matched internal comparison on the locked Flash, device-resident FP16 path. Matched Qwen2-7B full-stack prefill workload, sequence length 128, batch 16, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, measured in prompt positions per second. Internal, same-GPU, prefill-stage evidence.

~1.18x
Of the tested vLLM throughput at batch 16
Matched Qwen2-7B prefill workload, batch 16
~12%
Lower board joules per prefill position than the tested vLLM
NVML GPU-board energy, matched comparison
<1%
Throughput variation across measured runs
Stability of the matched comparison
Pass / fail†
Determinism contract: dual-run agreement on the measured Luxi path
What repeats is run-to-run reported behavior - not a receipt or output-token identity claim
What this is. Internal, prefill-stage evidence on a same-GPU matched comparison at batch 16, scoped to the tested configuration. Batch-32 data appears in the evidence pack as supporting detail. No matched B72 vLLM arm yet - this ratio is never attached to the batch-72 absolute result below. The independent TESTfort prefill baseline above is a separate, earlier evaluation and stays separate. Independent reproduction is the next step for this result.
Scope: Qwen2-7B-Instruct FP16, locked Flash device-resident path · sequence length 128 · batch 16 · one NVIDIA H100 80GB-class GPU · sequential comparison arms · matched prefill positions · GPU-board energy via NVML (see the reading guide above). The public brief linked above is the diligence artifact; the independent TESTfort evaluation (July 23, 2026) is preserved separately above with its third-party attribution.

Internal matched comparison, August 7, 2026. Recipe: locked Flash, device-resident FP16 path; shape: Qwen2-7B full-stack prefill, sequence length 128, batch 16; energy: NVML GPU-board energy. Reported separately from the independent TESTfort baseline above, which is a separate, earlier measurement of a different build. Version tags in the build lineage are internal milestone lineage, not products.

Matched evidence pack →

Luxi Inference Engine · Prefill · Internal Absolute Evidence

Internal absolute prefill champion - dual_gemm @ B72

The fastest measured configuration of the prefill path to date, reported as an absolute internal result with no vLLM comparison arm attached. One NVIDIA H100 80GB HBM3; Qwen2-7B-Instruct-class weights; sequence length 128; batch 72; dual-GEMM; Flash attention; device-resident FP16 path; median of five 15-second runs.

~44,860 /s
Prefill positions per second
Median of five 15-second runs
~0.01532 J
Board joules per prefill position
NVML GPU-board energy
What this is. Internal, absolute, prefill-stage evidence at batch 72. It is not a matched vLLM comparison: no matched B72 vLLM arm exists yet, so the ~1.18x batch-16 ratio above is never applied to this result.
Receipt result: the Door-B receipt check passed for independent-process stack-fingerprint equality under the dual-GEMM environment at batches 1, 16, and 32, with zero maximum difference under that declared contract. The B72 throughput loop itself did not emit hashes; its separate performance runs remained on Flash with no fallback. This is not a Door-A AUDIT bit-exact B72 result and not a universal determinism claim.

Internal absolute champion, August 7, 2026. Board energy via NVML (see the reading guide above). Prefill positions are not generated tokens; no decode or serving claim is made from this pack.

Absolute champion evidence pack →

Luxi Inference Engine · Decode/Generation · Internal Matched Evidence

Faithful full-model CUDA correctness

The current product story: LuxiEdge's faithful full-model Qwen2-7B CUDA path on NVIDIA H100, measured against CPU and Hugging Face references. Correctness contract: under the defined acceptance protocol, what repeats is the greedy token sequence - a pass/fail gate, not a performance claim.

Full 28-layer
Qwen2-7B on NVIDIA H100
Faithful full-model CUDA path
1.0**
CPU/CUDA greedy-token match rate
Current acceptance protocol
1.0**
HF/CUDA greedy-token match rate
Current acceptance test
~1.0
Logit cosine similarity
Top-10 overlap 10/10

Current status

  • Full 28-layer Qwen2-7B CUDA path on NVIDIA H100: running
  • CPU-versus-CUDA greedy-token match rate: 1.0 under the current acceptance protocol
  • Hugging Face-versus-CUDA greedy-token match rate: 1.0 under the current test
  • Logit cosine similarity: approximately 1.0 · top-10 overlap 10/10
  • Four-request independent CUDA generation smoke test: passed
  • All decoder mathematics runs on the GPU
  • Current limitations: token embedding row gathering and final greedy argmax remain host-side
  • Faithful-path performance, energy optimization, and production-serving validation: ongoing
Separation of claims: correctness results on this faithful path are separate from the TRADE energy benchmark in the research class below - energy and throughput numbers are never carried across the two.

**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.

Luxi Inference Engine · Decode/Generation · Internal Matched Evidence

Llama 3.1 internal milestone - independent validation next

Faithful Llama 3.1 inference - internal milestone

Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent validation of performance and energy is the next step; this section is separate from the public TRADE benchmark table.

Metric Internal result Scope
Correctness Active batch = serial Exact on primary p32 and p128 gates
Cold process/model load ~190–210 seconds Current development startup path
Status: internal Llama 3.1 resident-inference milestone with correctness gates passed. Independent validation of performance and energy is the next step. (Internal benchmark, not an MLPerf submission; energy is GPU-board energy per the reading guide above.)

Correctness gates passed internally; independent reproduction is on the validation roadmap.

Luxi Inference Engine · Decode/Generation · Research Evidence

Internal prototype · numerical audit complete

True-autoregressive decode - reference agreement achieved

Decode/generation produces output tokens. The GPU path was brought into reference agreement and matches the reference next-token result under the defined numerical audit. The next measurement milestone is post-audit end-to-end quality, throughput, and board energy on this path.

Provenance, August 6, 2026: true-AR throughput and board-joules figures published before this date are superseded and are not current product proof. The GPU route was rebuilt to the full reference mathematics, and the completed audit confirms it matches the reference next-token result.

Internal research evidence. Current true-autoregressive throughput and energy figures will be published once measured on this path; this entry is reported separately from the independent prefill evaluation, the later internal prefill measurement, and the fixed-window evaluated-position research.

Luxi Inference Engine · Decode/Generation · Research Evidence

Internal Research · data pack in preparation

Fixed-window verifier research - evaluated positions

An evaluated-position/verifier primitive on the decode research path. Work unit: evaluated positions per second - never generated tokens.

~4,669
Evaluated positions per second
Internal measurement · owner-supplied · data pack in preparation
Internal research. This is a verifier-side primitive measured in evaluated positions, distinct from serving or decode-throughput results. The public data pack is in preparation and this entry will link it when published.

Luxi Inference Engine · Research archive · Research Evidence

Internal TRADE research vs Hugging Face FP16

Separate internal controlled benchmark: TRADE research stack vs Hugging Face Transformers FP16 reference. This is not a production-peer comparison - it uses a different execution path and a different baseline than the matched-vLLM measurements above (the independent July 23, 2026 evaluation and the later internal measurement).

Up to 86%*
Lower board energy per processed token
~86.3% reduction measured
3.56×*
Processed-stack throughput
~32,083 vs ~9,026 stack-tokens/sec
~145 W
LuxiEdge median board power
Reference: ~298 W
H100 NVL
Qwen2-7B-Instruct · seq 128
Sustained-load measurement
MetricLuxiEdge TRADE research stackHugging Face Transformers FP16 referenceResult
Board energy per processed token~0.00453 J/token~0.03296 J/token~86.3% reduction
Processed-stack throughput~32,083 stack-tokens/sec~9,026 stack-tokens/sec~3.56×
Median board power~145 W~298 W~51% lower

NVIDIA H100 NVL · Qwen2-7B-Instruct · sequence length 128 · sustained-load measurement · controlled internal H100 research benchmark.

Research archive scope: these results measure the LuxiEdge TRADE research path under its stated sustain protocol - a research execution path with its own baseline, distinct from faithful-chat equivalence, vLLM comparisons, and production-serving evidence. For the current matched-vLLM measurement, see the internal prefill measurement section above.

*Measured as NVIDIA H100 board joules per processed token in a controlled sustained Qwen2-7B-Instruct research workload at sequence length 128. LuxiEdge TRADE research stack: approximately 0.00453 J/token; Hugging Face Transformers FP16 reference: approximately 0.03296 J/token, an approximately 86.3% reduction. The same test recorded approximately 3.56× processed-stack throughput. Internal benchmark; results vary by workload and system configuration. Faithful full-model Qwen2-7B CUDA integration is complete at the current correctness-acceptance level; faithful-path energy optimization and production-serving validation are ongoing.

Research Evidence

Long-context research

Memory scaling for the selectable long-context research path.

Sequence lengthWaller state (MB)Dense scores (MB)Reduction
1,0240.528.416×
4,0962.113464×
8,1924.2537128×
32,76816.84,295256×

Memory scales near-linearly vs dense O(N²). Pack →

Research Evidence · historical

Historical research configurations

Historical research archive, preserved with full methods and raw traces. For LuxiEdge's current claims and product state, see the sections above.

Earlier public benchmarks (LuxiDemo, 2026-07-11)

Earlier public evidence from a separate TRADE sustain protocol. Open any pack to inspect methods and raw traces.

7B-class TRADE - energy & throughput

Hardware: 1× NVIDIA H100 NVL · 28-layer stack · public pack.

Sequence length Tokens / sec Joules / token Median power (W) Notes
5~443.560 ± 0.005~156Short-burst reference
32~221--Throughput sweep
64~359--Throughput sweep
128~4030.630 ± 0.002~254Ladder primary (this protocol)
256~464~0.604~280Longer context
512~493--Throughput sweep

Multi-run at sequence length 128, from the earlier public sustain ladder - a different protocol than the Qwen2-7B benchmark above; the two are not directly comparable. Source: evidence/h100-7b-class-TRADE/

Head-to-head with a Flash baseline

12 layers · sequence 1024 · h=768 · ~10 s each side.

MetricTRADE 12LPyTorch + Flash 12LRatio
Median power (W)177.2176.4~1.0×
ms / stack forward74.323.90TRADE 19.1× slower
Prefill tok/s13778262859PT 19.1× higher
Joules / token0.01290.0007TRADE 19.2× higher

h100-stack12-H2H - we publish where a standard baseline wins on this shape.

Evidence packs

Open on GitHub. Read methods. Check traces.

Absolute champion

prefill_accel_lock_20260807T233111Z

Internal absolute prefill champion: dual-GEMM at batch 72, ~44,860 positions/s at ~0.01532 board J/position, median of five 15-second runs. Absolute internal result; no matched vLLM arm.

Matched comparison

prefill_freeze_matched_20260807T210749Z

Internal matched comparison at batch 16: ~1.18x the tested vLLM throughput at ~12% lower board J/position, with batch-32 supporting data. Qwen2-7B-Instruct FP16, seq 128, sequential comparison arms.

Independent baseline

h100-qwen2-7b-v99-matched-prefill-2026-07-23

Independent TESTfort prefill baseline (July 23, 2026): ~28,375 positions/s at ~0.0187 board J/position - 80.60% of default vLLM throughput at 3.10% lower board energy per position.

Milestone lineage

version-100-h100-gtm

Earlier internal matched prefill evidence from the internal milestone lineage (version tags are internal milestones, not products). Superseded as the current headline by the August 7, 2026 packs above.

Earlier protocol

h100-7b-class-TRADE

Full 28-layer 7B-class TRADE stack from the earlier public sustain ladder. Multi-run throughput + energy (~0.63 J/tok @ seq=128, ~403 tok/s under that historical protocol).

Energy

h100-stack12-TRADE-cuda

Device-resident 12-layer stack energy (GPT-2 width).

Honesty

h100-stack12-H2H

TRADE vs PyTorch Flash - complete head-to-head, both sides published.

AUDIT

h100-WNSM-free-ride

Null-space free-ride under GPU load.

Memory

h100-LONGCTX-scaling

O(N) vs O(N²) memory ladder + CUDA 32k.

Micro

h100-BASELINE-vs-geo

Single-layer baseline wedges - not a 7B throughput claim.

Serve

h100-serve-sustain-2026-07-11

Continuous-batch serve sustain with power traces.

Open LuxiDemo on GitHub

Download

Public demo builds for the expression/receipt path. Demo only - not production inference. Energy claims live in the packs above, not only inside these binaries.

Open latest release →

LuxiEdge Lite / Demo

macOS ARM, Linux x86 (CPU/GPU), Linux ARM, Windows - from the public release channel.

Quick start

chmod +x luxiedge-*-macos-arm64
./luxiedge-*-macos-arm64 --port 9090

curl -X POST http://localhost:9090/evaluate \
  -H "Content-Type: application/json" \
  -d '{"expr":"sin(x)","values":[0.5,1.0],"precision":"f32"}'
Demo binaries may expire. Full operator sets require a license - contact.

Two different kinds of hashes appear around these builds: release-file checksums (published alongside downloads) verify that the file you downloaded was not corrupted or altered in transit; evaluation-output receipts (SHA-256 hashes emitted by the engine) attest what a specific computation produced under a documented deterministic mode. One is about file integrity, the other about computation results - they are not interchangeable.