Luxi Inference Engine · Prefill stage · controlled evaluation
The Inference Prefill Evaluation, measured on your terms.
The Luxi Inference Engine runs transformer workloads in two stages: prefill, which reads the prompt, and decode/generation, which produces output tokens. Prefill is the current evaluation offer - a paid or funded engagement that measures the Luxi prefill stage against your current serving stack on a prompt-heavy workload, with the measurement contract agreed in writing before anything runs. This evaluation covers the prefill stage; decode/generation is a later stage and not part of this offer.
What we agree before anything runs
The evaluation is defined by a written measurement contract. Both sides sign off on every item below before the first run.
Useful-work unit
The agreed unit of work - for example prompt positions processed. Prompt positions are not generated decode tokens, and the report will never present them as such.
Model and shapes
The exact model, precision, sequence lengths, and batch sizes to be measured. Results apply to those shapes only.
Hardware
The GPU class and host environment for both arms of the comparison - for example one NVIDIA H100 80GB-class GPU per arm.
Reference runtime
Your current stack, pinned: for example a specific vLLM version and configuration. LuxiEdge is measured against that reference - we do not claim LuxiEdge replaces vLLM as a serving stack.
Determinism contract
What must repeat, exactly: which outputs, under which rerun protocol, and how agreement is checked.
Throughput requirement and energy boundary
The throughput your workload must satisfy, and the energy boundary that will be reported - GPU-board energy via NVML, defined precisely in the measurement contract.
Reference evidence - two separate classes
These are the public reference points for what a prefill evaluation measures. They are different lanes with different provenance, and we do not merge them.
Independent prefill baseline - third-party, 2026-07-23
TESTfort third-party packed-prefill measurement (batch 16): independently verified 3.10% lower board J/position vs default vLLM at 80.60% of default vLLM throughput. Preserved as dated independent evidence with the complete measured tradeoff published.†
Baseline detail → · TestFort validation report - July 2026 (PDF) →
Internal prefill evidence - absolute and matched, reported separately
Internal absolute prefill: roughly 44,860 prefill positions per second at roughly 0.01532 board joules per position - one H100 80GB HBM3, Qwen2-7B-Instruct-class weights, sequence length 128, batch 72, dual-GEMM, Flash attention, device-resident FP16, median of five 15-second runs; no matched vLLM arm. Internal matched comparison: at batch 16, roughly 1.18x the tested vLLM throughput at roughly 12% lower board joules per position, with throughput variation below 1% across measured runs. Both internal, same-GPU, prefill-stage evidence, reported separately from each other and from the independent baseline; independent reproduction is the next step.‡
Absolute champion → · Public data brief → · Matched comparison →
†The independently tested build carries the internal milestone tag Version 99. ‡The internal build carries the internal milestone tag Version 100. These tags are internal milestones, not product names.
Exactly what you are buying
Engagement shape
This is a paid or funded evaluation. Custom integration and workload analysis begin under a paid evaluation, funded design partnership, investment arrangement, or strategic agreement. Public evidence stays public; customer-specific measurement is scoped work, handled under an NDA where appropriate.
- Read the public evidence: the prefill evidence and the public data brief.
- Tell us your workload, reference runtime, hardware, and throughput requirement.
- We draft the measurement contract together, then run the comparison.