Scoped Evaluation Open

Product ยท measured comparison on agreed terms

A scoped H100 prefill evaluation, measured on your terms.

A paid or funded engagement that measures LuxiEdge against your current serving stack on a prompt-heavy workload - with the measurement contract agreed in writing before anything runs.

What we agree before anything runs

The evaluation is defined by a written measurement contract. Both sides sign off on every item below before the first run.

Useful-work unit

The agreed unit of work - for example prompt positions processed. Prompt positions are not generated decode tokens, and the report will never present them as such.

Model and shapes

The exact model, precision, sequence lengths, and batch sizes to be measured. Results apply to those shapes only.

Hardware

The GPU class and host environment for both arms of the comparison - for example one NVIDIA H100 80GB-class GPU per arm.

Reference runtime

Your current stack, pinned: for example a specific vLLM version and configuration. LuxiEdge is measured against that reference - we do not claim LuxiEdge replaces vLLM as a serving stack.

Determinism contract

What must repeat, exactly: which outputs, under which rerun protocol, and how agreement is checked.

Throughput requirement and energy boundary

The throughput your workload must satisfy, and the energy boundary that will be reported - GPU board energy via NVML, not facility, wall-plug, or PUE-adjusted energy.

Reference evidence - two separate classes

These are the public reference points for what an evaluation measures. They are different lanes with different provenance, and we do not merge them.

Version 100 - internal measurement, July 2026

Matched-vLLM prefill comparison on one H100 80GB-class GPU, Qwen2-7B-Instruct FP16, sequence length 128, batches 16 and 32: about 1.17-1.18× higher prompt-position throughput and about 10-14% lower board joules per position than the tested vLLM stack. Internal measurement - not a TESTfort or third-party evaluation.

Public brief → · Full table →

Version 99 - third-party, 2026-07-23

TESTfort third-party packed-prefill measurement (batch 16): 3.10% lower board J/position vs default vLLM at 80.60% of default vLLM throughput. Preserved as dated earlier evidence - including the fact that LuxiEdge did not win throughput on that workload.

Version 99 detail → · TestFort July 2026 report (PDF) →

What this evaluation is not

Boundaries, stated plainly. This is not a claim that LuxiEdge replaces vLLM or any production serving engine. It is not a multi-tenant serving-stack comparison and not an every-workload claim. Prompt positions per second are not full chat decode tokens per second. Board energy via NVML is not facility, wall-plug, or PUE-adjusted energy. The report measures the agreed configuration - nothing beyond it.

Engagement shape

This is a paid or funded evaluation. Custom integration and workload analysis begin under a paid evaluation, funded design partnership, investment arrangement, or strategic agreement. Public evidence stays public; customer-specific measurement is scoped work, handled under an NDA where appropriate.

  1. Read the public evidence: Proof and the Version 100 brief.
  2. Tell us your workload, reference runtime, hardware, and throughput requirement.
  3. We draft the measurement contract together, then run the comparison.