Our mission: more useful AI from every watt.
Building AI inference that uses less energy.
LuxiEdge is designed to pair deterministic execution with measurable efficiency improvements, helping data centers get more useful work from the power they already have.
The Version 100 H100 benchmark below is evidence of that mission - not the entire mission. Its claims apply only to the specifically tested Qwen2-7B-Instruct FP16 prefill workload and the tested vLLM configuration.
~1.17-1.18× throughput · ~10-14% lower board J/position vs vLLM†
Batch 16 and 32 · Qwen2-7B-Instruct FP16 · H100 · Version 100
Faithful full-model Qwen2-7B CUDA execution now running on NVIDIA H100, with 1.0 token agreement in the current Hugging Face acceptance test.
†Version 100 internal measurement, July 2026. Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32, one NVIDIA H100 80GB-class GPU, sequential comparison arms, matched prefill positions, one generated token per iteration. Batch 16: LuxiEdge ~41,700 pos/sec, ~0.0171 J/pos vs vLLM ~35,800 pos/sec, ~0.0190 J/pos (~1.17× throughput, ~10% lower board energy). Batch 32: LuxiEdge ~43,900 pos/sec, ~0.0158 J/pos vs vLLM ~37,100 pos/sec, ~0.0182 J/pos (~1.18× throughput, ~13% lower board energy). Prompt positions per second are not full chat decode tokens per second. Board energy via NVML - not facility, wall-plug, or PUE-adjusted energy. This is not a multi-tenant serving-stack or every-workload claim. This measurement is not a TESTfort or third-party evaluation. Public brief: H2H Prefill Energy Brief.
**Current acceptance testing includes matched CPU, CUDA, and Hugging Face greedy-token comparisons under the defined limited protocol. It is a correctness result, not a broad model-quality or performance ranking.
Full methods and downloads: Proof.
Measurement scope: one NVIDIA H100 80GB-class GPU · Qwen2-7B-Instruct FP16 · batches 16 and 32 · sequence length 128 · matched prefill positions · one generated token per iteration · NVML GPU-board energy. Prompt positions per second are not full chat decode tokens per second. Not facility, wall-plug, or PUE-adjusted energy. Not a multi-tenant serving-stack or every-workload result. Detailed methods →
Faithful Llama 3.1 resident inference
Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent performance and energy validation pending.
Correctness gates passed internally. Independent reproduction and performance/energy validation are pending. Not an official MLPerf result.
A three-legged stool
Pick the path that fits you. Every path can end in shareable proof.
Quant & research
Determinism, audit trails, free-ride paths, long-context memory - for people who need methods, not slogans.
AI & data centers
Power caps, density, measured joules-per-token under load, and a commercial path to evaluation.
We need your help
You don’t need a PhD. Share the energy story. Ask AI providers and data centers why they aren’t cutting waste.