JevBench v1.5.4 · individual system

Instinct Dual 4B

Closed decision API · by ZooWork / pierre-srp · Marked closed

Base model: Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.

v1.5 roster addendum A4

Not independently recorded

JevBench v1.5.4 score

46.968

Option A: rank #27 of 106 ranked systems.

The three option scores and ranks are published independently; the headline is Option A.

Published JevBench option scores and ranks
OptionScoreRank
A · headline46.968#27 of 106
B42.974#31 of 106
C32.617#32 of 106

Published axes

intelligence
43.1
calibration
88.3
speed
82.0
cost
60.2

Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.

Run and cost evidence

Run status
complete · 1,624 decisions · 0 missing
Cost
$0.021 per 1,000 decisions · estimate · ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.
Median latency
0.506 seconds, adjusted · x2 (demo assumption, not measured)
Endpoint condition
operator-hosted free-preview API
API exposure
The endpoint received sealed item text · API measurement: the operator's endpoint received sealed item text, without answers.
Published source
https://github.com/fstandhartinger/jevbench/issues/123
Model and serving disclosure

Operator-hosted free-preview demo measured 27/28 Sep 2026. No immutable model artifact is externally verifiable; the pin is endpoint/model ID instinct-dual-4b on the measurement date. Operator reports Qwen3.5-4B, two option-order probability passes, no weight training and zero generated tokens; these architecture claims are self-reported. The author disclosed selecting the method using all 231 published items. ZooWork had already received all 720 sealed item texts in an earlier evaluation; no labels were supplied. Costs use the frozen exact-base reference estimate, not the announced $0.01/M tariff. Latency uses the frozen demo adjustment (2x).

Values come from the public v1.5.4 aggregate. Scores and ranks may change in a later release.

Read the full leaderboard, the v1.5.4 release page, and the published method.