JevBench v1.5.4 · individual system

Nemotron Diffusion 8B (pst2154, optimized vLLM)

system-one-open · by pst2154 · Code and weights marked open

Base model: nvidia/Nemotron-Labs-Diffusion-8Bsource

v1.5 roster addendum A3

MIT (application code); modified vLLM runtime Apache-2.0; nvidia/Nemotron-Labs-Diffusion-8B keeps its upstream licence

JevBench v1.5.4 score

25.685

Option A: rank #45 of 106 ranked systems.

The three option scores and ranks are published independently; the headline is Option A.

Published JevBench option scores and ranks
OptionScoreRank
A · headline25.685#45 of 106
B22.704#47 of 106
C17.837#55 of 106

Published axes

intelligence
33.8
calibration
76.2
speed
93.7
cost
55.5

Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.

Run and cost evidence

Run status
complete · 1,624 decisions · 0 missing
Cost
$0.030 per 1,000 decisions · estimate · documented hosted-model estimate; no exact base-model floor applies
Median latency
0.187 seconds, adjusted · x2 + 0.15 s (assumption, not measured)
Endpoint condition
evaluator-owned Lium GPU pod (H100), offline read-only container
Published source
https://github.com/pst2154/Nemotron_Jev
Model and serving disclosure

nvidia/Nemotron-Labs-Diffusion-8B @ 16c67f05 (unchanged weights) served by github.com/pst2154/Nemotron_Jev @ cdc36d60 on the author's digest-pinned modified vLLM runtime: BF16, candidate-only output projection, direct candidate logits, CUDA graphs, one masked forward pass per decision, nothing generated. Disclosed by the author: the inherited option-ordering rule was selected using public benchmark examples in earlier 14B experiments. Our earlier v1.4.3 run of this system used a recipe that left the base image's older 14B application in place; this v1.5 row is a fresh measurement of the 8B.

Values come from the public v1.5.4 aggregate. Scores and ranks may change in a later release.

Read the full leaderboard, the v1.5.4 release page, and the published method.