AI & data centers
AI data-center power constraints and inference energy evidence.
The story for operators is three separate and clearly labeled paths: the internal matched prefill measurement (matched-vLLM, internal, July 2026) is the current production-peer energy and throughput result; separate TRADE research vs Hugging Face FP16 provides an additional internal energy benchmark; and the faithful CUDA path establishes scoped model correctness. None of these results are interchangeable. Honest numbers, not throughput-only marketing.
LuxiEdge for AI Data Centers
Measured GPU board energy · Null-space transport research · Separate reproducibility path
The research-lane measurements behind this sketch are archived with their full protocol, comparators, and warnings on Proof. GPU board energy measured with NVML - not facility or wall-plug energy. Results apply to the tested configuration. Methodology and evidence: Proof →
Internal prefill evidence - absolute and matched, reported separately
Two internal results on the locked Flash, device-resident FP16 path, one NVIDIA H100 80GB-class GPU, Qwen2-7B-Instruct FP16, sequence length 128, measured in prefill positions per second. Internal, prefill-stage evidence; independent reproduction is the next step.
| Batch | vs vLLM throughput | vs vLLM board energy |
|---|---|---|
| 16 | ~1.18x the tested vLLM | ~12% lower board J/position |
| 72 | No matched B72 vLLM arm yet | Absolute only: ~44,860 positions/s at ~0.01532 J/position |
Recipe: locked Flash, device-resident FP16 path; shape: Qwen2-7B full-stack prefill, seq 128, batch 16. Energy: GPU-board energy via NVML. Results apply to the tested workload and configuration. Full methods: Proof →
Faithful Llama 3.1 inference - internal milestone
Internal Llama 3.1 resident-inference milestone. Active-batch output matched the serial BF16 path on tested correctness gates. Independent validation of performance and energy is the next step.
Correctness gates passed internally; independent reproduction is on the validation roadmap.
Two pillars (not a throughput-only pitch)
1. Energy under load
- Joules per token you can open in a public pack
- Sustained work - not launch-tax microbenches as the hero
- Power while the stack is actually running
- Same method language the public uses to pressure providers
2. Determinism & audit
- Receipt-oriented paths where trust matters
- Free-ride / residual checks under load
- Reproducible story for multi-tenant and regulated AI
- Methods you can re-run - not “it looked about the same”
Where speed fits (honestly)
LuxiEdge is not “ignore throughput.” The controlled Qwen2-7B research benchmark recorded higher processed-stack throughput than the tested reference (archived with its protocol at Proof) - and we still publish where standard baselines win. The positioning is:
What we optimize for
Energy cost of work + deterministic trust, with throughput that holds up under the protocols we publish. Long-context memory scaling (O(N) vs O(N²)) when context is the silent capacity killer.
What you can count on
Every number here was measured on a stated workload with the comparator named, and the complete head-to-head is published - both sides, including the 12-layer Flash comparison. Differentiated axes: audit paths, memory scaling, and full-stack energy you can verify.
Why operators evaluate us
Power caps are real
Rack and site power - not TFLOPS slides - decide how much AI you can sell. J/token is how you plan.
Trust is multi-tenant
Deterministic paths and receipts reduce “it drifted” incidents across customers and re-runs.
Public pressure is coming
Communities and customers are asking about AI electricity. See We need your help - the same numbers work for operators and the public.
Proof before NDA
Headline metrics link to public packs. Public evidence and available source are provided through the project’s public repositories; private technical diligence, customer-specific materials, and confidential integration work may be provided under an executed NDA when appropriate.
Next step
Data-center operators, inference providers, and infrastructure partners are invited to evaluate LuxiEdge against their own workloads and systems.
- Read Proof - energy tables and evidence first.
- Map your power cap and throughput SLO to the published protocols.
- Technical discussion for on-prem / cloud evaluation and joint metering.
- Ready to engage? See the scoped H100 prefill evaluation offer.
Help Make AI Energy Measurable
Most people do not know who operates the GPUs behind the AI they use. Share this evidence with your AI service, cloud provider, employer's infrastructure team, sustainability team, or anyone responsible for purchasing and operating AI systems.