Hosting = where inference runs; company = where the provider or lab is registered.
EU hosting means the route's inference runs inside the EU: an EU region, AWS Bedrock's EU cross-region (geo) profiles, Azure's Europe Data Zone, or a provider whose entire public fleet is documented as EU-hosted, each checked per model against the provider's documentation. Global deployments do not count, and neither does an EU billing region, an EU company or an EU control plane on its own. One disclosed company-policy exception stays in, marked “EU equivalent”.
Data confidentialitywhat the provider may do with your prompts
More settings
Evidence requirements (benchmark evidence, priced provider, measured task tokens) sit above the table they apply to.
Applies to price views & model offers; benchmark evidence stays unfiltered.
JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
Frozen release · 1,624 decisions per system (904 open + 720 sealed; sealed decisions are half of Intelligence) · 105 ranked of 111 roster systems · only system-level sealed aggregates are published · aggregate results JSON sha256 5c3d97440ebb…
$0.011 per 1,000 decisions = 0.35× Jev → green: as cheap as Jev or better
Latency vs cap
0.28 s = 0.46× Jev → green: as fast as Jev or better
Rank
Capability Score 59 · official #80
Colours
Green ≤ reference; amber ≤ cap; red > cap. Shorter is cheaper or faster.
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev, ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = 0.07, n = 105).
Jev (TypeSafe, closed)
Jev rebuild
Instruction model, JSON schema
Small tool-calling model
Zero-shot classifier
Closed decision API
Reranker (neutral adapter)
Raw-logit control (base model)
Native-logit decision engine
Unclassified
system-one-open
green: ≤ reference
amber: ≤ cap (2× reference)
red: > cap
Outside the Jev-class limits · 46 systems
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev 1.13.0 ($0.032 per 1,000 decisions, median 0.62 s).
$0.019 per 1,000 decisions = 0.59× Jev → green: as cheap as Jev or better
Latency vs cap
58.16 s = 94.34× Jev → red: outside the 2× cap
Rank
Capability Score outside Jev-class · official #97
Colours
Green ≤ reference; amber ≤ cap; red > cap. Shorter is cheaper or faster.
Jev-class = cost per decision at most 2× Jev 1.13.0's (≤ $0.065 per 1,000 decisions)and median latency at most 2× Jev 1.13.0's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 59 of 105 systems qualify; the other 46, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed
Jev-class systems are shown by default. Bubble size follows the official JevBench Score. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
59 systems. Hover or focus a bubbleTap a bubble for its values.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
59 systems. Hover or focus a bubbleTap a bubble for its values.
Jev (TypeSafe, closed)
Jev rebuild
Instruction model, JSON schema
Small tool-calling model
Zero-shot classifier
Closed decision API
Reranker (neutral adapter)
Raw-logit control (base model)
Native-logit decision engine
Unclassified
system-one-open
faint = outside Jev-class
What-If: the weight sliders below re-score every system under other axis weights — only the equal 25/25/25/25 weights give the official option-A ranking. The 3D view of capability, cost and speed loads further down.
JevBench v1.5.3
JevBench Composite Score: 105 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).
Whiskers are 95% bootstrap intervals. 69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other 27 adjacent pairs, so no tie classification is inferred.
All 105 systems
Greener = stronger within its column.
105 of 105 systems, sorted by official rank, #1 first.
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; new = first listed in v1.5.3; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Click a column heading to sort. Names link to each project.
Honorable mentions and Jev wrappers — listed separately, not ranked (1)
Eligibility rule: services that run on Jev itself may be measured and shown as honorable mentions, but are not competitors ranked against Jev and do not enter the field median gap (G_med) or tie markers. classifier.dev (TypeSafe) runs on Jev, so it stays unranked under this rule.
classifier.dev (fast tier)APIhonorable mention: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2). Official (A) score 74.7.
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
0–100, the values in the table. A system with no published axis draws at 0 and says so.
Capability by subject topic
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.What each category means · items per category
Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 235 items (174 open / 61 sealed)
Coding & software — code, SQL, repositories, developer tools and IT systems. 92 items (67 open / 25 sealed)
Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 548 items (260 open / 288 sealed)
Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 282 items (131 open / 151 sealed)
Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 115 items (98 open / 17 sealed)
Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 128 items (40 open / 88 sealed)
Use cases (TypeSafe categories)
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.What each category means · items per category
Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 54 items (18 open / 36 sealed)
Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 132 items (131 open / 1 sealed)
Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 98 items (61 open / 37 sealed)
Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 389 items (176 open / 213 sealed)
Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 31 items (10 open / 21 sealed)
Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 508 items (261 open / 247 sealed)
E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 71 items (48 open / 23 sealed)
Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 78 items (39 open / 39 sealed)
Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 15 items (11 open / 4 sealed)
Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 159 items (76 open / 83 sealed) — not a use case of its own, so it is counted but not drawn.
Low n (under 15 items, not plotted): Scientific discovery (1), Semantic code linting (4), Feature extraction for predictive modeling (7), Recruiting (3), Moderation and trust and safety (12), Advertising (2), Gaming (0), Demand forecasting (0).
Competence per request type, open / sealed
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 904 open and 720 sealed decisions.
Competence per tier — open set
Per-tier competence, the three request types pooled by their published decision counts.
Competence per tier — sealed set
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made
Chance-corrected competence (0 = chance, 100 = perfect) over the category's items, open and sealed pooled, all tiers: per request type the method's competence cell (Choice and Noul accuracy above chance, Score 1 - error / chance error), then averaged over types weighted by item count. Not part of the JevBench Score; compare systems within a category, not categories with each other. Values can be negative (below chance).
Every one of the 1,624 decisions was labelled with one subject topic (the v1.2 topic list, unchanged) and one or two TypeSafe use-case categories (docs.typesafe.ai/concepts/use-case-map, plus Other) by Winnow-12B Q8 (the model behind System1 Models s1-pro) on our own GPU pod; 1,118 of the 3,248 calls were first made on s1-pro itself and agree 98.2 % with the pod run. Sealed items stayed on our own infrastructure. About 5 % of public items were checked by hand (Claude Opus 5.5); the item-group rules below fix the systematic misses found.
routing, routing_hard and tool_selection items (choose the model, agent or tool that handles a request) are Model routing first; the labeller tagged many by the request subject.
tool_guardrail, unsafe_action, adequacy, judge_hard and policy_compliance items (check an agent tool call or a drafted answer before it goes out) are LLM guardrails first.
lead_qualification items take the topic Finance & commerce (sales); the labeller read the scoring rubric as rules & law.
A second use case is kept when the labeller gave it probability >= 0.15 and it is not Other.
All values as a table
Spoke
A: Jev 1.13.0
B: Cygnet
The four score axes
Intelligence
72.0
71.1
Calibration
88.0
87.0
Speed
83.8
91.0
Cost
54.7
56.4
Capability by subject topic
Math & numbers
40.5
31.4
Coding & software
80.3
87.8
Rules, policy & law
73.7
76.2
Finance & commerce
83.4
78.5
Support & operations
91.5
88.2
Everyday language
96.1
95.6
Safety & security
67.7
74.4
Use cases (TypeSafe categories)
Search & retrieval
95.9
94.7
Model routing
96.5
99.1
LLM guardrails
67.0
72.3
Lead generation
88.1
82.2
Customer support
91.9
88.9
Insurance claims
67.5
71.7
Financial crime
58.3
71.1
Legal & compliance
74.0
75.3
E-commerce
89.8
91.6
Risk assessment
80.1
73.1
Knowledge graphs
100.0
73.2
Competence per request type, open / sealed
Choice · open
85.7
82.9
Choice · sealed
87.6
80.3
Noul · open
47.8
57.9
Noul · sealed
48.6
55.7
Score · open
81.2
76.6
Score · sealed
81.1
73.2
Competence per tier — open set
Easy
90.0
89.7
Standard
80.9
82.4
Judge
83.5
82.2
Hard
53.8
58.0
Competence per tier — sealed set
Easy
91.2
88.6
Standard
79.0
73.5
Judge
73.7
75.4
Hard
72.9
65.3
Axes, request types, latency and cost
Compare:
Every measured system, under the four score axes. Intelligence is 50% open (904 decisions) and 50% sealed (720); Gap = I_open − I_sealed, and the penalty applies only above the field median gap (G_med 5.2) plus 8. The per-type competence, latency and price of the same systems are in Types & cost. On a phone the name column stays put while the table scrolls sideways.
The same systems, by what was asked and what it cost. Per-type columns are chance-corrected competence (CC, 0 = chance) for Choice, Noul and Score, open / sealed; I open and I sealed are the two halves of Intelligence. Latency is adjusted p50 / p95; cost is per 1,000 decisions. On a phone the name column stays put while the table scrolls sideways.
API = the operator's endpoint received sealed item text during evaluation, without answers. Sealed item text, answers and item-level results stay private; only system-level aggregates appear here. Hover a cost for its price basis and a latency for its raw values and adjustment.
Official order with 95% intervals (105 systems)
JevBench v1.5.3 · headline option A
JevBench Score: 105 ranked systems
Official (A)weighted harmonic mean of four 0–100 axes, Intelligence · Calibration · Speed · Cost = 25 · 25 · 25 · 25, with the low-axis gates · Method ↓
Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).
Whiskers are 95% bootstrap intervals. 69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other 27 adjacent pairs, so no tie classification is inferred.
1CygnetBase model:google/gemma-4-12B-itsource73.7I 71 · C 87 · S 91 · K 56 · B#2 · C#1 · $0.028
2Winnow-12B Q8Base model:google/gemma-4-12B-itsource73.2I 74 · C 84 · S 86 · K 57 · B#1 · C#2 · $0.028
3Jev 1.13.0APIBase model:undisclosed72.1I 72 · C 88 · S 84 · K 55 · B#3 · C#3 · $0.032
4JevK5 v0.3v1.5 roster addendum A1Base model:Qwen/Qwen3.5-4Bsource71.9I 56 · C 88 · S 94 · K 63 · B#5 · C#8 · $0.017
5Plumb-4Bv1.5 roster addendum A1Base model:alibiserikbay/JevK5source71.6I 56 · C 87 · S 93 · K 63 · B#6 · C#9 · $0.017
6Jev-OmniBase model:google/gemma-4-12B-itsource71.5I 70 · C 83 · S 85 · K 56 · B#4 · C#4 · $0.029
7decider-4b v2Base model:Qwen/Qwen3.5-4B-Basesource71.3I 56 · C 86 · S 91 · K 65 · B#7 · C#10 · $0.015
8Decision 4B v1.2v1.5 roster addendum A1Base model:undisclosed70.8I 54 · C 89 · S 94 · K 63 · B#9 · C#12 · $0.017
9Imajev-4Bv1.5 roster addendum A2Base model:Qwen/Qwen3.5-4Bsource70.4I 53 · C 88 · S 91 · K 63 · B#11 · C#13 · $0.017
10Decision 4B v1.1v1.5 roster addendum A1Base model:undisclosed70.4I 53 · C 87 · S 94 · K 63 · B#12 · C#14 · $0.017
11SemIfBase model:Qwen/Qwen3.5-4Bsource68.7I 51 · C 84 · S 91 · K 63 · B#14 · C#19 · $0.017
12spark-s1-4b-v6Base model:Qwen3.5-4Bsource68.2I 62 · C 70 · S 86 · K 60 · B#8 · C#5 · $0.021
13metask-jev-4bBase model:Qwen3.5-4Bsource67.5I 53 · C 83 · S 89 · K 58 · B#15 · C#16 · $0.026
14HopperBase model:Qwen/Qwen3.5-4Bsource67.5I 50 · C 88 · S 87 · K 62 · B#18 · C#23 · $0.018
15Malkuth-4BBase model:Qwen/Qwen3.5-4B-Basesource66.8I 54 · C 83 · S 86 · K 56 · B#16 · C#15 · $0.030
16Surogate Rune 26B-A4B v3v1.5 roster addendum A2Base model:google/gemma-4-26B-A4B-itsource66.5I 70 · C 88 · S 86 · K 49 · B#10 · C#6 · $0.050
17reflex 4BBase model:Qwen/Qwen3.5-4Bsource65.2I 51 · C 87 · S 69 · K 63 · B#19 · C#20 · $0.017
18jev-localBase model:Qwen/Qwen3.5-9Bsource65.2I 56 · C 78 · S 73 · K 59 · B#17 · C#11 · $0.024
19djev (Maisa, diffusion-gemma)Base model:google/diffusiongemma-26B-A4B-itsource64.2I 72 · C 80 · S 91 · K 48 · B#13 · C#7 · $0.053
20Raw Qwen3 4B Instruct 2507 direct logitsBase model:undisclosedsource62.1I 54 · C 53 · S 89 · K 63 · B#20 · C#18 · $0.017
21jqvBase model:Qwen3-32Bsource60.8I 49 · C 87 · S 83 · K 51 · B#21 · C#25 · $0.042
22JevK5 v0.2.0Base model:Qwen3.5-4Bsource58.1I 47 · C 85 · S 91 · K 63 · B#22 · C#28 · $0.017
23Qwen3.5-9B Jev-like data-mix v2Base model:Qwen/Qwen3.5-9Bsource53.0I 60 · C 81 · S 82 · K 46 · B#23 · C#17 · $0.065
24Standard One 8BBase model:mistralai/Ministral-3-8B-Instruct-2512-BF16source47.8I 60 · C 83 · S 92 · K 43 · B#26 · C#22 · $0.078
25NInfer Qwen3.8-Flash-Next mixedBase model:Qwen3.8-Flash-Nextsource47.5I 67 · C 89 · S 89 · K 43 · B#24 · C#21 · $0.082
26Instinct Dual 4Bv1.5 roster addendum A4APIBase model:Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.47.0I 43 · C 88 · S 82 · K 60 · B#30 · C#31 · $0.021
27swanOneBase model:Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source46.6I 71 · C 87 · S 85 · K 42 · B#25 · C#24 · $0.085
28Raw Qwen3 8B direct logitsBase model:Qwen/Qwen3-8B-Basesource45.2I 51 · C 49 · S 87 · K 46 · B#27 · C#30 · $0.065
29decider-2bBase model:Qwen/Qwen3.5-2B-Basesource45.1I 42 · C 72 · S 94 · K 65 · B#33 · C#33 · $0.015
30system-oneBase model:Qwen3-8Bsource44.1I 50 · C 49 · S 91 · K 45 · B#28 · C#34 · $0.068
31system-one-openAPIBase model:Gemma 4 E2Bsource42.4I 42 · C 72 · S 78 · K 68 · B#34 · C#36 · $0.011
32Autoloops – Gemma 4 31B ITAPIBase model:undisclosed40.5I 77 · C 86 · S 84 · K 40 · B#31 · C#26 · tariff$0.103
33GPT-6 Luna (low reasoning effort)APIBase model:undisclosed40.5I 95 · C 95 · S 73 · K 39 · B#29 · C#27 · $0.108
34GPT-6 Luna (default medium reasoning effort)APIBase model:undisclosed38.8I 96 · C 96 · S 73 · K 38 · B#32 · C#29 · $0.114
35JevOneBase model:Qwen/Qwen3.6-35B-A3Bsource38.2I 54 · C 85 · S 90 · K 40 · B#35 · C#35 · $0.101
36kev 4BBase model:Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.38.1I 40 · C 68 · S 85 · K 66 · B#36 · C#39 · $0.014
37kev 8BBase model:Qwen/Qwen3-8B-Basesource34.2I 48 · C 71 · S 84 · K 40 · B#37 · C#41 · $0.097
38open-alternative-jevBase model:Qwen3.5-4Bsource33.6I 37 · C 77 · S 91 · K 63 · B#39 · C#42 · $0.017
39Bespoke Nimble 9BBase model:Qwen3.5-9Bsource31.8I 64 · C 77 · S 83 · K 37 · B#38 · C#32 · $0.128
40Malkuth-2BBase model:empero-ai/Qwen3.8-2B-Distillsource29.9I 36 · C 75 · S 92 · K 66 · B#42 · C#44 · $0.014
41openjev-sglangAPIBase model:Qwen/Qwen3.6-35B-A3Bsource29.0I 59 · C 83 · S 78 · K 36 · B#40 · C#37 · $0.140
42decider-35b-a3bBase model:Qwen/Qwen3.5-35B-A3B-Basesource27.5I 60 · C 82 · S 91 · K 34 · B#41 · C#38 · $0.154
43local-jev Qwen3.5-4BBase model:Qwen/Qwen3.5-4Bsource25.8I 34 · C 82 · S 84 · K 59 · B#45 · C#53 · $0.023
44Nemotron Diffusion 8Bv1.5 roster addendum A3Base model:nvidia/Nemotron-Labs-Diffusion-8Bsource25.7I 34 · C 76 · S 94 · K 55 · B#46 · C#54 · $0.030
45Open-Jev 9BBase model:Qwen/Qwen3.5-9Bsource24.4I 64 · C 82 · S 74 · K 33 · B#43 · C#40 · $0.170
46Decision 2BBase model:openbmb/MiniCPM5-2Bsource22.5I 31 · C 86 · S 90 · K 66 · B#52 · C#56 · $0.013
47GPT-5.6 LunaAPIBase model:undisclosed22.4I 94 · C 95 · S 74 · K 31 · B#44 · C#43 · tariff$0.205
48typecastlmBase model:Qwen3.5-4Bsource21.8I 31 · C 77 · S 92 · K 64 · B#56 · C#58 · $0.016
49JEV Qwen3.5-9B Base NVFP4Base model:ig1/Qwen3.5-9B-NVFP4source20.1I 32 · C 81 · S 94 · K 48 · B#58 · C#59 · $0.056
50Gemini 3.1 Flash-LiteAPIBase model:undisclosed19.6I 78 · C 75 · S 80 · K 30 · B#47 · C#45 · tariff$0.219
51AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1Base model:Qwen/Qwen3.8-27Bsource19.5I 73 · C 88 · S 88 · K 29 · B#48 · C#46 · $0.226
52AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2Base model:Qwen/Qwen3.8-27Bsource19.5I 73 · C 87 · S 87 · K 29 · B#49 · C#47 · $0.226
53NInfer Qwen3.8-27B NVFP4Base model:Qwen3.8-27Bsource18.7I 65 · C 86 · S 90 · K 29 · B#51 · C#48 · $0.231
54Eikos-27Bv1.5 roster addendum A1Base model:Qwen/Qwen3.8-27Bsource18.5I 75 · C 86 · S 88 · K 29 · B#50 · C#49 · $0.238
55NInfer Qwen3.8-27B NVFP4 (T=1.5)Base model:Qwen3.8-27Bsource18.5I 61 · C 86 · S 90 · K 29 · B#54 · C#50 · $0.231
56InstinctAPIBase model:Qwen3.8-27Bsource · Operator-reported; weights not publicly verifiable.18.3I 63 · C 85 · S 82 · K 29 · B#55 · C#51 · $0.230
57OpenJev (thinking, BF16)Base model:undisclosedsource17.9I 84 · C 83 · S 74 · K 29 · B#53 · C#52 · $0.241
58djev (thinking)Base model:google/diffusiongemma-26B-A4B-itsource17.4I 77 · C 96 · S 72 · K 28 · B#57 · C#55 · $0.249
59LitJevBase model:Qwen/Qwen3.8-27Bsource16.3I 58 · C 84 · S 68 · K 28 · B#59 · C#57 · $0.244
60Bev / Bonsai 27Bv1.5 roster addendum A4Base model:Qwen/Qwen3.8-27Bsource15.8I 53 · C 78 · S 73 · K 28 · B#60 · C#61 · $0.247
61Raw Phi-4 mini direct logitsBase model:undisclosedsource15.2I 28 · C 71 · S 89 · K 53 · B#62 · C#62 · $0.036
62OpenSourceJevBase model:Qwen3.5-4Bsource13.2I 26 · C 76 · S 73 · K 69 · B#63 · C#63 · $0.011
63reflex-27bBase model:Qwen3.8-27Bsource13.2I 63 · C 86 · S 69 · K 26 · B#61 · C#60 · $0.297
64Open-Jev 2BBase model:Qwen/Qwen3.5-2Bsource9.1I 34 · C 74 · S 76 · K 33 · B#64 · C#65 · $0.170
65GLiNER2 largeBase model:microsoft/deberta-v3-largesource8.4I 22 · C 42 · S 65 · K 78 · B#66 · C#66 · $0.0056
66Qwen3-Reranker-4BBase model:Qwen/Qwen3-4B-Basesource7.1I 21 · C 76 · S 80 · K 48 · B#67 · C#67 · $0.052
67DeepSeek V4.1 FlashAPIBase model:undisclosed6.6I 94 · C 97 · S 69 · K 19 · B#65 · C#64 · $0.498
68SimpleJevBase model:Qwen/Qwen3.5-0.8Bsource4.0I 17 · C 47 · S 59 · K 70 · B#69 · C#69 · $0.0098
69SimpleJev Qwen3.8-27BAPIBase model:Qwen/Qwen3.8-27Bsource3.4I 73 · C 87 · S 75 · K 15 · B#68 · C#68 · $0.687
70decision-machine-1APIBase model:undisclosed3.2I 15 · C 81 · S 93 · K 56 · B#70 · C#70 · $0.029
71GLiNER2.5 multiBase model:mDeBERTa-v3-basesource2.7I 14 · C 58 · S 67 · K 87 · B#71 · C#71 · $0.0028
72Bosun v3.1 0.6Bv1.5 roster addendum A4Base model:Qwen/Qwen3-0.6Bsource2.5I 14 · C 65 · S 62 · K 78 · B#72 · C#72 · $0.0056
73GLiNER2Base model:DeBERTa-v3-basesource2.3I 13 · C 36 · S 70 · K 87 · B#73 · C#73 · $0.0028
74Deem 0.8B v1v1.5 roster addendum A4Base model:Qwen/Qwen3.5-0.8Bsource2.1I 13 · C 39 · S 84 · K 80 · B#74 · C#74 · $0.0045
75JevActAPIBase model:undisclosedsource1.5I 11 · C 63 · S 76 · K 69 · B#76 · C#75 · $0.011
76CLM-8BBase model:Qwen/Qwen3-8Bsource1.5I 11 · C 48 · S 93 · K 51 · B#75 · C#76 · $0.045
77kev 0.6BBase model:Qwen/Qwen3-0.6B-Basesource1.3I 10 · C 68 · S 87 · K 80 · B#77 · C#77 · $0.0046
78Raw Qwen3 0.6B direct logitsBase model:Qwen/Qwen3-0.6B-Basesource1.1I 11 · C 21 · S 90 · K 78 · B#78 · C#78 · $0.0056
79GLiNER2.5 smallBase model:microsoft/deberta-v3-xsmallsource0.9I 9 · C 56 · S 77 · K 87 · B#80 · C#79 · $0.0028
80Raw Qwen3 1.7B direct logitsBase model:Qwen/Qwen3-1.7B-Basesource0.9I 10 · C 22 · S 90 · K 69 · B#79 · C#80 · $0.011
81MirrorBase model:undisclosed0.2I 6 · C 43 · S 64 · K 89 · B#81 · C#81 · $0.0023
82ZeroEntropy zerank-2Base model:Qwen/Qwen3-4Bsource0.1I 5 · C 82 · S 80 · K 48 · B#82 · C#82 · $0.052
83jeffBase model:undisclosedsource0.1I 4 · C 80 · S 56 · K 81 · B#83 · C#83 · $0.0043
84smalljev semantic-v9Base model:openbmb/MiniCPM5-2B-Basesource0.1I 4 · C 73 · S 86 · K 61 · B#84 · C#84 · $0.020
85Laya multilingualv1.5 roster addendum A4Base model:mmBERT-basesource0.0I 2 · C 44 · S 74 · K 82 · B#85 · C#85 · $0.0039
86OpenDecisionBase model:MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.0I 2 · C 73 · S 87 · K 79 · B#86 · C#86 · $0.0050
87BAAI bge-reranker-v2-m3Base model:BAAI/bge-m3source0.0I 0 · C 83 · S 91 · K 59 · B#87 · C#87 · $0.023
88Certo v1Base model:ModernBERT-largesource0.0I 0 · C 88 · S 91 · K 97 · B#88 · C#88 · $0.0013
89Decision FastBase model:Qwen/Qwen3-0.6B-Basesource0.0I 0 · C 76 · S 91 · K 80 · B#89 · C#89 · $0.0046
90Alibaba GTE Reranker ModernBERT-baseBase model:answerdotai/ModernBERT-basesource0.0I 0 · C 75 · S 91 · K 69 · B#90 · C#90 · $0.011
91kev 0.5BBase model:Qwen/Qwen2.5-0.5Bsource0.0I 0 · C 65 · S 88 · K 80 · B#91 · C#91 · $0.0046
92LayaBase model:ModernBERT-largesource0.0I 0 · C 74 · S 74 · K 85 · B#92 · C#92 · $0.0032
93lev-350mBase model:LiquidAI/LFM2.5-350Msource0.0I 0 · C 78 · S 94 · K 80 · B#93 · C#93 · $0.0046
94Qwen3.5-0.8B Decision ModelBase model:Qwen/Qwen3.5-0.8B-Basesource0.0I 0 · C 74 · S 72 · K 80 · B#94 · C#94 · $0.0048
95Mixedbread mxbai-rerank-base-v2Base model:undisclosedsource0.0I 0 · C 87 · S 89 · K 60 · B#95 · C#95 · $0.021
96Needle 3Base model:undisclosedsource0.0I 0 · C 0 · S 34 · K 62 · B#96 · C#96 · $0.019
97Needle 3, options as toolsBase model:undisclosedsource0.0I 0 · C 0 · S 41 · K 62 · B#97 · C#97 · $0.019
98open-jev-deberta-v3-largeBase model:microsoft/deberta-v3-largesource0.0I 0 · C 77 · S 68 · K 78 · B#98 · C#98 · $0.0056
99Open Jev JSON CanvasBase model:google/diffusiongemma-26B-A4B-itsource0.0I 77 · C 0 · S 86 · K 49 · B#99 · C#99 · $0.049
100openJev VerdictBase model:undisclosedsource0.0I 0 · C 52 · S 84 · K 87 · B#100 · C#100 · $0.0028
101openJev Verdict 1.4Base model:knowledgator/gliclass-modern-base-v2.0source0.0I 0 · C 80 · S 81 · K 87 · B#101 · C#101 · $0.0028
102Qwen3.8 27BAPIBase model:undisclosed0.0I 96 · C 98 · S 57 · K 0 · B#102 · C#102 · $2.18
103verdict-smallBase model:intfloat/multilingual-e5-smallsource0.0I 0 · C 59 · S 82 · K 100 · B#103 · C#103 · $0.0009
104VonBase model:answerdotai/ModernBERT-largesource0.0I 0 · C 83 · S 76 · K 83 · B#104 · C#104 · $0.0038
105Laya typed-decisionsv1.5 roster addendum A4Base model:ModernBERT-largesource0.0I 0 · C 83 · S 63 · K 83 · B#105 · C#105 · $0.0038
Costs are estimates (est.) unless marked tariff.
system-one-openJev rebuildJev (TypeSafe, closed)UnclassifiedRaw-logit control (base model)Native-logit decision engineClosed decision APIInstruction model, JSON schemaZero-shot classifierReranker (neutral adapter)Small tool-calling model
All three weight options (105 systems)
All three weight options
A is the official headline: equal 25/25/25/25 axis weights and an Intelligence floor of 50. B remains the secondary 40/20/20/20 axis-weight view; C retains equal axes with an Intelligence floor of 60. All three use equal Choice/Noul/Score weights. The CI column is the paired-bootstrap 95% interval of the A score.
Among the 59 Jev-class systems, Jev 1.13.0 has the highest JevBench Capability Score, 80.0 (Intelligence 72.0, Calibration 88.0).
Cygnet leads the official JevBench v1.5.3 score (option A) with 73.7: Intelligence 71.1, Calibration 87.0, Speed 91.0, Cost 56.4 ($0.028 per 1,000 decisions).
The best open or open-planned rebuild, Winnow-12B Q8, is #2 at 73.2 — 0.5 points behind.
GPT-6 Luna (default medium reasoning effort) has the highest Intelligence (96.2) but places #34: Speed 73.2, Cost 38.3 — the harmonic mean does not let accuracy buy back a weak axis.
The strongest sealed Intelligence is 98.2 (Qwen3.8 27B, #102); sealed items carry half of Intelligence, and an open-minus-sealed gap beyond the field median plus eight points costs Intelligence.
69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other adjacent pairs, so no tie classification is inferred.
All 16 systems joined by separately hashed roster addenda and have official ranks in this revision; their A/B/C ranks and intervals are in the addendum table below.
classifier.dev scores 74.7 but is not ranked: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2).
SimpleJev Qwen3.6-35B-A3B, Decision-4B did not complete the full suite; they are listed without a rank.
Jev alternatives, open source and self-hosting
The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.
What are open-source alternatives to Jev?
The highest-ranked open entrants in this run are Cygnet (#1, 73.7), Winnow-12B Q8 (#2, 73.2), JevK5 v0.3 (#4, 71.9), Plumb-4B (#5, 71.6). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.
Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?
Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.
jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.
How is JevBench scored?
The official score (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost — 25/25/25/25 — with an Intelligence floor of 50 and low-axis gates on Speed and Cost. Version v1.5.3 measures 904 open and 720 sealed decisions per system; Choice, Noul and Score each carry a third, sealed items contribute 50% of Intelligence, and an open-minus-sealed gap beyond the field median costs points. Method notes · options B and C.
How do I submit my model?
Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version or a disclosed roster addendum. For private data, see the custom evaluation options.
What a decision costs
Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole typed request — state, rubric and options — not a single token.
How costs are estimated
Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured — 3 rows carry a tariff. Systems without one — open weights, author demos, models we ran ourselves — are priced as if a large inference provider hosted them: the list price of the same weights, or the nearest larger sibling or size class when the exact weights are not listed. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.
Price rules (v1.5). Only public, bookable list prices that have been in effect for at least 30 days count; a manufacturer's standard, non-promotional launch list price counts from day one, and promotions, subsidies, credits and free tiers never do. The scoring price is never below the market reference price of the system's base model. A system without any eligible price is listed as unpriced — no Cost axis and no score until a price qualifies. A later price change triggers a re-score with a visible note on the row.
API models with a known base model (from 1 Oct 2026). They are ranked at their developer's own stated API price; a striped second bar shows the score and rank they would have at the base-model reference price, the way we price self-served open weights of the same base. First applied on Image JevBench v0.1.5 (Wity-1). The v1.5.4 scores on this page are unchanged; JevBench rows follow this rule from the next version.
Cygnet — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Winnow-12B Q8 — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Jev 1.13.0 — ~$0.032 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
Jev-Omni — ~$0.029 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
decider-4b v2 — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
SemIf — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
spark-s1-4b-v6 — ~$0.021 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
metask-jev-4b — ~$0.026 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Hopper — ~$0.018 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Malkuth-4B — ~$0.030 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
reflex 4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
jev-local — ~$0.024 est. per 1,000 decisions: price floor: base-model reference price applied
djev (Maisa, diffusion-gemma) — ~$0.053 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Raw Qwen3 4B Instruct 2507 direct logits — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
jqv — ~$0.042 est. per 1,000 decisions: documented hosted-model estimate
JevK5 v0.2.0 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Qwen3.5-9B Jev-like data-mix v2 — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
Standard One 8B — ~$0.078 est. per 1,000 decisions: documented hosted-model estimate
NInfer Qwen3.8-Flash-Next mixed — ~$0.082 est. per 1,000 decisions: documented hosted-model estimate
swanOne — ~$0.085 est. per 1,000 decisions: price floor: base-model reference price applied
Raw Qwen3 8B direct logits — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
decider-2b — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate
system-one — ~$0.068 est. per 1,000 decisions: documented hosted-model estimate
system-one-open — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
GPT-6 Luna (low reasoning effort) — ~$0.108 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
GPT-6 Luna (default medium reasoning effort) — ~$0.114 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
JevOne — ~$0.101 est. per 1,000 decisions: documented hosted-model estimate
kev 4B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
kev 8B — ~$0.097 est. per 1,000 decisions: price floor: base-model reference price applied
open-alternative-jev — ~$0.017 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Bespoke Nimble 9B — ~$0.128 est. per 1,000 decisions: documented hosted-model estimate
Malkuth-2B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
openjev-sglang — ~$0.140 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
decider-35b-a3b — ~$0.154 est. per 1,000 decisions: price floor: base-model reference price applied
local-jev Qwen3.5-4B — ~$0.023 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Open-Jev 9B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
Decision 2B — ~$0.013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
typecastlm — ~$0.016 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
JEV Qwen3.5-9B Base NVFP4 — ~$0.056 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference
NInfer Qwen3.8-27B NVFP4 — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
NInfer Qwen3.8-27B NVFP4 (T=1.5) — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
Instinct — ~$0.230 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
OpenJev (thinking, BF16) — ~$0.241 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
djev (thinking) — ~$0.249 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
LitJev — ~$0.244 est. per 1,000 decisions: price floor: base-model reference price applied
Raw Phi-4 mini direct logits — ~$0.036 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
OpenSourceJev — ~$0.011 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
reflex-27b — ~$0.297 est. per 1,000 decisions: price floor: base-model reference price applied
Open-Jev 2B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
GLiNER2 large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
Qwen3-Reranker-4B — ~$0.052 est. per 1,000 decisions: ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
DeepSeek V4.1 Flash — ~$0.498 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
SimpleJev — ~$0.0098 est. per 1,000 decisions: documented hosted-model estimate
SimpleJev Qwen3.8-27B — ~$0.687 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
decision-machine-1 — ~$0.029 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
GLiNER2.5 multi — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
GLiNER2 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
JevAct — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
CLM-8B — ~$0.045 est. per 1,000 decisions: price floor: base-model reference price applied
kev 0.6B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Raw Qwen3 0.6B direct logits — ~$0.0056 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
GLiNER2.5 small — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
Raw Qwen3 1.7B direct logits — ~$0.011 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Mirror — ~$0.0023 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
ZeroEntropy zerank-2 — ~$0.052 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
jeff — ~$0.0043 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
smalljev semantic-v9 — ~$0.020 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
OpenDecision — ~$0.0050 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
BAAI bge-reranker-v2-m3 — ~$0.023 est. per 1,000 decisions: ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3)
Certo v1 — ~$0.0013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Decision Fast — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Alibaba GTE Reranker ModernBERT-base — ~$0.011 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies
kev 0.5B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Laya — ~$0.0032 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
lev-350m — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Qwen3.5-0.8B Decision Model — ~$0.0048 est. per 1,000 decisions: documented hosted-model estimate
Mixedbread mxbai-rerank-base-v2 — ~$0.021 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies
Needle 3 — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
Needle 3, options as tools — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
open-jev-deberta-v3-large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
Open Jev JSON Canvas — ~$0.049 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
openJev Verdict — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
openJev Verdict 1.4 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
Qwen3.8 27B — ~$2.18 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
verdict-small — ~$0.0009 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Von — ~$0.0038 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
classifier.dev — ~$0.023 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies
JevK5 v0.3 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Plumb-4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Decision 4B v1.2 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Imajev-4B — ~$0.017 est. per 1,000 decisions: ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B.
Decision 4B v1.1 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
Surogate Rune 26B-A4B v3 — ~$0.050 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
AutoJev-27B (denis-pplx, Qwen3.8-27B) — ~$0.226 est. per 1,000 decisions: documented hosted-model estimate
AutoJev-27B (RTX PRO 6000) — ~$0.226 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
Eikos-27B — ~$0.238 est. per 1,000 decisions: documented hosted-model estimate
SimpleJev Qwen3.6-35B-A3B — ~$0.145 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
Nemotron Diffusion 8B — ~$0.030 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
Bev / Bonsai 27B — ~$0.247 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b.
Deem 0.8B v1 — ~$0.0045 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B.
Decision-4B — ~$0.014 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.
Instinct Dual 4B — ~$0.021 est. per 1,000 decisions: ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.
Bosun v3.1 0.6B — ~$0.0056 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate.
Laya multilingual — ~$0.0039 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.
Laya typed-decisions — ~$0.0038 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.
Roster addendum: newcomers scored on the same frozen protocol (16)
All 16 A1/A2/A3/A4 systems below completed all 1,624 decisions and now have official ranks in A, B, and C; the interactive score presets also include them. Their scores, individual 95% intervals, and the frozen v1.5.0 G_med are unchanged. No new paired-bootstrap comparisons were computed for addendum systems. Existing tie markers are retained only for base-system pairs that remain adjacent; no tie or separation is inferred for the other pairs.
Ranks follow option scores; score intervals are per system, not pairwise rank comparisons. Every slider preset sorts these same ranked systems using its selected weights.
Not ranked: partial, unpriced and unmeasured systems
These systems are part of the 111-system v1.5 roster but have no rank. Their numbers are never shown as zero or free.
Partial runs (2)
SimpleJev Qwen3.6-35B-A3BAPIpartial run: Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.
Decision-4B (Eval Engine / Chromia)v1.5 roster addendum A4not ranked: PARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.
Incomplete or not measured in v1.5 (3)
Jobe Qwen3.5-4B (frozen): not measured. Base model:undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.
mica-v01-4bv1.5 roster addendum A1: The frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank. (1,088/1,624 rows; 536 missing). Base model:undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16): not measured. Base model:undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.
Incomplete and unmeasured systems receive no official rank. Existing results from earlier benchmark versions remain on their frozen version pages.
Method notes: what changed in v1.5
The four axes
Intelligence · 25%
How often answers are right above chance: each type is normalized against its task-specific random baseline (chance = 0, perfect = 100; below-chance tiers can be negative). Choice, Noul and Score count one third each. Easy, Standard, Judge and Hard items count 10%, 20%, 30% and 40%; the open and sealed sets count equally.
Calibration · 25%
How closely stated probabilities match what happens. It uses ECE and TVD for Choice, ECE and Brier for Noul, and normalized RPS plus top-level ECE for Score. The three types count equally; open and sealed items are pooled.
Speed · 25%
Serial response latency on open Standard and Judge items. The p50 and p95 each get a log score: 100 − 20 × log₁₀(seconds ÷ 0.1), then are averaged. Self-hosted and demo endpoints get the published ×2 plus 0.15-second adjustment.
Cost · 25%
Estimated or billed US dollars per 1,000 decisions, using pooled token use across 1,624 decisions and the documented price rules. The log score is 100 − 30 × log₁₀(cost ÷ $0.001). The price reference is $0.001 per 1,000 decisions.
The official score is the weighted harmonic mean of the four axes. Option A gives each axis 25%; option B (40/20/20/20) is a secondary view, while option C keeps equal axes and sets the Intelligence gate at 60. In A and B, Intelligence, Speed and Cost each have a quadratic gate below 50. To limit benchmaxxing on the public items, the open-minus-sealed Intelligence gap may be up to 8 points above the field median (G_med) before a penalty applies. Each further point lowers the multiplier on unpenalized Intelligence by one percentage point.
The method owner chose equal axis weights and equal weights for Choice, Noul and Score after reviewing the What-If Lab, preserving continuity with v1.4 and treating the three decision types equally. Disclosed headline amendment: equal-axis, equal-type A, SHA-256 752ddccc4e19…. B remains a secondary view.
Re-evaluation policy. Every release re-evaluates the current top 10 on the composite score. Models ranked #11 and below are re-evaluated on a slower cadence — at least monthly, or with every third scheduled refresh release, whichever comes first — and their score is shown as last measured on its release. A material method change re-evaluates every model. Paid fast-lane runs are evaluated within 48 hours of payment, and new submissions are evaluated in the order received.
1,624 decisions per system: 904 open (601 published) and 720 sealed, drawn fresh from a private pool with the same tier mix as the open set. Sealed counts for 50% of Intelligence: base = 0.5 × I_open + 0.5 × I_sealed.
Three request types are scored natively and chance-corrected per item: Choice, Noul and Score each receive one third. Tier weights easy / standard / judge / hard = 10 / 20 / 30 / 40. A type a system does not support is excluded, never scored zero; only full-coverage systems are ranked.
Overfit penalty relative to the field: excess = gap − G_med, penalty = max(0, 1 − max(0, excess − 8) / 100). G_med for this batch is 5.2 CC points.
Calibration is typed (Choice ECE/TVD, Noul ECE with Brier, Score normalised RPS and top-level ECE), pooled over open and sealed. Speed and Cost formulas are unchanged from v1.4; self-hosted and demo endpoints carry the ×2 + 0.15 s adjustment. A manufacturer's standard, non-promotional launch list price counts from day one, but a newer price cut younger than 30 days does not. Rows without token counts use the measured proxy-token basis. A system without any eligible public, bookable price is listed as unpriced.
The frozen 25 Sep DeepInfra snapshot records Qwen3.5-4B as deprecated on 11 Jun 2026 and replaced by Qwen3.5-9B. Its frozen snapshot rates remain the v1.5 M2 reference; price basis tooltips and the correction note disclose this. Pricing disclosure correction SHA-256: 1b660648bd49….
Composite: weighted harmonic mean with the Intelligence, Speed and Cost gates below 50 (Intelligence below 60 in option C). The official headline A uses equal 25 / 25 / 25 / 25 axis weights and Intelligence floor 50. B remains the secondary 40 / 20 / 20 / 20 view; C keeps equal axes and Intelligence floor 60. Ties come from the paired bootstrap. The gates belong to the score, not to the weights: in the custom-weight views of the score chart they still apply to an axis set to 0, as in the published views — so “Intelligence only” follows the Intelligence column except for systems with Cost, Speed or Intelligence below 50 (a general-purpose LLM with Cost 39 keeps only (39/50)², about 0.61, of its score). Every gated row carries a “gate ×…” tag naming the axis and factor, and the weights panel lists the gates that fire.
All 16 full-coverage systems marked with a v1.5 roster addendum label are officially ranked in A, B, C, and every score preset. Their scores, individual intervals, method, pricing rules and frozen G_med are unchanged. No new paired-bootstrap comparison is inferred for addendum pairs.
Before every release we review the leaderboard for anomalies and close loopholes with general, documented rules. The page and Git repository provide transparent data and method details; Benchmark Heaven owns its rules.
Data file SHA-256 5c3d97440ebb1463133ce1049df78cf3fd829c5f735c053c92cf707e7b3f7b24 · scorer output SHA-256 452885de2a84cd5b9ed393d541fd8d6c6a540f9d7f762f9d2df9349383f26173 · run kind official.
Limits
1,624 decisions per system (904 open, 720 sealed) is a measurement, not a census, and it is English-only.
The weights are a choice. Option A weights the four axes equally and uses a harmonic mean, so the weakest axis dominates; options B and C are published alternatives and the weight sliders re-score the same axes for exploration — only the official option gives the official score and rank. If a wrong decision costs you more than a slow or expensive one, read the Intelligence column and the per-type competence rather than the score alone.
The latency adjustment (×2, +0.15 s on our own servers and demo endpoints) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. Serving under load trades per-user speed for throughput. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo.
Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
Latency is one origin at one time of day; hosted endpoints, public demos and our own pods are different kinds of latency. Public demo endpoints are shared with everyone else using them.
Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit
Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.
Decision 2B (FlyMy.AI, v59) — FlyMy.AI (@denti), Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial release — huggingface.co/flymy-ai/decision-2b-preview
Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA) — unknown, not recorded
Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA) — unknown, not recorded
Decision Fast (FlyMy.AI, v53a) — FlyMy.AI (@denti), Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial release — huggingface.co/flymy-ai/decision-fast-preview
GPT-5.6 Luna (low reasoning effort) — OpenAI, proprietary API
GPT-6 Luna (default medium reasoning effort) — OpenAI, proprietary API
GPT-6 Luna (low reasoning effort) — OpenAI, proprietary API
Hopper — HopitAI, Component-specific terms recorded in RESULT.md; submitted adapter release and Qwen base retain their respective terms — huggingface.co/HopitAI/hopper
jev-local (Qwen3.5-9B) — us (GitHub), no licence stated in the repository (public code); Apache-2.0 base weights — github.com/us/jev-local
Jev-Omni (akhilaaa3, Gemma-4-12B merged) — akhilaaa3, Apache-2.0, following Gemma 4; dataset rights stated separately by the author — huggingface.co/akhilaaa3/Jev-Omni
Malkuth-2B (newfull5, Kev post-train) — newfull5 (dhtocks), CC-BY-NC-4.0, research use only (XNLI and RACE in the training mix) — github.com/newfull5/malkuth
Malkuth-4B (newfull5, Kev post-train) — newfull5 (dhtocks), CC-BY-NC-4.0, research use only (XNLI and RACE in the training mix) — github.com/newfull5/malkuth
Open-Jev 2B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-Jev
Open-Jev 9B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-Jev
openjev-sglang (Qwen3.6-35B-A3B on SGLang) — ekzhang, no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms — github.com/ekzhang/openjev-sglang
OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp) — sabeel111, MIT (repository code); Apache-2.0 (Qwen/Qwen3.5-4B base and unsloth/Qwen3.5-4B-GGUF Q4_K_M conversion) — github.com/sabeel111/OpenSourceJev
Plumb-4B (crh225, JevK5 v0.2 + LoRA) — unknown, not recorded
spark-s1-4b-v6 (Open Spark Jev, abhishek085) — Abhishek Rai (abhishek085), Apache-2.0 (code and weights); base Qwen/Qwen3.5-4B Apache-2.0 — github.com/abhishek085/open-spark-jev
Standard One 8B (Standard Thinking) — Standard Thinking (myeongho12), Apache-2.0 (jev-adapter server, LoRA and merged weights; base Ministral 3 8B Apache-2.0) — huggingface.co/StandardThinking/StandardOne-8B
Surogate Rune 26B-A4B v3 (RTX PRO 6000) — unknown, not recorded
swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4) — blockbrain, patches Apache-2.0, shim MIT; weights under the Qwen licence (LICENSE-NOTICE.md) — github.com/blockbrain-ai/swanone-recipe
system-one (Qwen3-8B, Sean Goedecke) — Sean Goedecke, no licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0 — github.com/sgoedecke/system-one
system-one-open (Gemma 4 E2B LoRA on an L4) — mithalouni, MIT (repository LICENSE; Gemma weights keep Google’s terms) — github.com/mithalouni/system-one-open
Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version or a disclosed roster addendum rather than silently changing this one.