JevBench v1.5.4 — Jev alternatives ranking

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Frozen release · 1,624 decisions per system (904 open + 720 sealed; sealed decisions are half of Intelligence) · 106 ranked of 112 roster systems · only system-level sealed aggregates are published · aggregate results JSON sha256 0cf210b76bf8…

Making decisions from images? Explore Image JevBench v0.1.5 and compare its systems.

Share this version · View live board · Previous release: JevBench v1.5.3

JevBench v1.5.4 · headline

JevBench Capability Score

Capability ranking of Jev-class systems

Capability Score averages Intelligence and Calibration. Jev 1.13.0 leads the Jev-class systems with 80.0.

Jev-class means at most 2× the cost and median latency of Jev. How we choose ↘

Adjust cost / latency caps · 2× official
Official
#SystemICCap.$/1k
  1. 1Jev 1.13.0API72.054.780.0$0.032*Reference system for the Jev-class limits
  2. 2Winnow-12B Q874.456.679.3$0.028*
  3. 3Cygnet71.156.479.0$0.028*
  4. 4Surogate Rune 26B-A4B v369.749.079.0$0.050*
  5. 5Jev-Omni70.556.176.5$0.029*
  6. 6djev72.348.276.4$0.053*
  7. 7JevK5 v0.356.363.172.3$0.017*
  8. 8Plumb-4B55.863.171.6$0.017*
  9. 9Decision 4B v1.253.763.171.1$0.017*
  10. 10Imajev-4B53.563.370.8$0.017*
Show all 60 Jev-class systems (50 more)
  1. 11decider-4b v255.864.570.7$0.015*
  2. 12Decision 4B v1.153.163.170.2$0.017*
  3. 13Hopper49.962.368.9$0.018*
  4. 14Malkuth-4B54.555.868.8$0.030*
  5. 15metask-jev-4b53.557.768.1$0.026*
  6. 16Manchego v2.151.264.368.1$0.015*
  7. 17jqv49.151.267.9$0.042*
  8. 18SemIf51.363.167.6$0.017*
  9. 19jev-local56.358.967.0$0.024*
  10. 20spark-s1-4b-v662.160.465.9$0.021*
  11. 21JevK5 v0.2.046.763.165.8$0.017*
  12. 22Instinct Dual 4BAPI43.160.265.7$0.021*
  13. 23Decision 2B31.366.458.8$0.013*
  14. 24local-jev Qwen3.5-4B33.759.358.1$0.023*
  15. 25open-alternative-jev37.463.357.1$0.017*
  16. 26decider-2b42.364.956.9$0.015*
  17. 27system-one-openAPI41.668.356.6$0.011*
  18. 28JEV Qwen3.5-9B Base NVFP432.247.656.6$0.056*
  19. 29Malkuth-2B35.565.655.4$0.014*
  20. 30Nemotron Diffusion 8B33.855.555.0$0.030*
  21. 31typecastlm31.264.054.0$0.016*
  22. 32kev 4B39.965.853.8$0.014*
  23. 33Raw Qwen3 4B Instruct 2507 direct logits54.163.453.5$0.017*
  24. 34Raw Phi-4 mini direct logits27.653.249.5$0.036*
  25. 35Qwen3-Reranker-4B21.048.448.6$0.052*
  26. 36decision-machine-1API14.856.348.0$0.029*
  27. 37Certo v10.096.943.8$0.0013*
  28. 38Mixedbread mxbai-rerank-base-v20.060.343.3$0.021*
  29. 39ZeroEntropy zerank-24.848.443.2$0.052*
  30. 40Von0.082.741.7$0.0038*
  31. 41BAAI bge-reranker-v2-m30.059.341.7$0.023*
  32. 42openJev Verdict 1.40.086.640.2$0.0028*
  33. 43kev 0.6B10.380.139.0$0.0046*
  34. 44lev-350m0.080.038.8$0.0046*
  35. 45Open Jev JSON Canvas77.149.338.6$0.049*
  36. 46smalljev semantic-v93.760.838.4$0.020*
  37. 47Decision Fast0.080.138.0$0.0046*
  38. 48Alibaba GTE Reranker ModernBERT-base0.069.337.6$0.011*
  39. 49OpenDecision2.179.137.3$0.0050*
  40. 50JevActAPI11.268.636.9$0.011*
  41. 51GLiNER2.5 small9.286.632.5$0.0028*
  42. 52kev 0.5B0.080.132.5$0.0046*
  43. 53CLM-8B11.450.530.0$0.045*
  44. 54verdict-small0.0100.029.7$0.00087*
  45. 55openJev Verdict0.086.626.1$0.0028*
  46. 56Deem 0.8B v113.080.326.0$0.0045*
  47. 57GLiNER213.586.624.5$0.0028*
  48. 58Laya multilingual2.482.223.0$0.0039*
  49. 59Raw Qwen3 0.6B direct logits10.677.616.0$0.0056*
  50. 60Raw Qwen3 1.7B direct logits9.768.515.6$0.011*

Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev, ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.

Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = 0.07, n = 106).

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
  • green: ≤ reference
  • amber: ≤ cap (2× reference)
  • red: > cap
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev 1.13.0 ($0.032 per 1,000 decisions, median 0.62 s).

  1. –Qwen3.8 27BAPI95.60.096.8$2.18*Outside: cost 67.4× Jev, latency 10.5× Jev
  2. –GPT-6 Luna (medium)API96.238.395.9$0.11*Outside: cost 3.5× Jev, latency 2.5× Jev
  3. –DeepSeek V4.1 FlashAPI93.719.195.3$0.50*Outside: cost 15.4× Jev, latency 2.9× Jev
  4. –GPT-6 Luna (low)API95.339.195.1$0.11*Outside: cost 3.3× Jev, latency 2.6× Jev
  5. –GPT-5.6 LunaAPI94.330.794.5$0.20Outside: cost 6.3× Jev, latency 2.1× Jev
  6. –djev77.328.186.5$0.25*Outside: cost 7.7× Jev, latency 2.4× Jev
  7. –OpenJev84.228.583.7$0.24*Outside: cost 7.5× Jev, latency 2.5× Jev
  8. –Autoloops – Gemma 4 31B ITAPI76.739.681.3$0.10Outside: cost 3.2× Jev
  9. –Eikos-27B75.128.780.7$0.24*Outside: cost 7.4× Jev
  10. –AutoJev-27B72.829.480.2$0.23*Outside: cost 7.0× Jev
  11. –SimpleJev Qwen3.8-27BAPI72.814.980.0$0.69*Outside: cost 21.3× Jev, latency 2.7× Jev
  12. –AutoJev-27B72.829.479.7$0.23*Outside: cost 7.0× Jev
  13. –swanOne71.242.279.1$0.085*Outside: cost 2.6× Jev
  14. –NInfer Qwen3.8-Flash-Next mixed67.242.577.9$0.082*Outside: cost 2.5× Jev
  15. –Gemini 3.1 Flash-LiteAPI77.629.876.1$0.22Outside: cost 6.8× Jev
  16. –NInfer Qwen3.8-27B NVFP465.529.175.7$0.23*Outside: cost 7.1× Jev
  17. –reflex-27b62.825.874.4$0.30*Outside: cost 9.2× Jev, latency 4.5× Jev
  18. –InstinctAPI62.729.273.9$0.23*Outside: cost 7.1× Jev
  19. –NInfer Qwen3.8-27B NVFP461.129.173.8$0.23*Outside: cost 7.1× Jev
  20. –Open-Jev 9B63.833.172.7$0.17*Outside: cost 5.3× Jev, latency 2.1× Jev
  21. –LitJev58.328.471.4$0.24*Outside: cost 7.6× Jev, latency 5.0× Jev
  22. –Standard One 8B59.643.371.4$0.078*Outside: cost 2.4× Jev
  23. –decider-35b-a3b60.534.471.2$0.15*Outside: cost 4.8× Jev
  24. –openjev-sglangAPI58.635.670.8$0.14*Outside: cost 4.3× Jev
  25. –Qwen3.5-9B Jev-like data-mix v260.445.770.7$0.065*Outside: cost 2.0× Jev
  26. –Bespoke Nimble 9B63.736.870.5$0.13*Outside: cost 4.0× Jev
  27. –JevOne54.139.869.5$0.10*Outside: cost 3.1× Jev
  28. –reflex 4B51.563.169.2$0.017*Outside: latency 4.6× Jev
  29. –Bev / Bonsai 27B53.128.265.4$0.25*Outside: cost 7.6× Jev, latency 3.6× Jev
  30. –kev 8B48.340.459.7$0.097*Outside: cost 3.0× Jev
  31. –Open-Jev 2B33.633.153.6$0.17*Outside: cost 5.3× Jev
  32. –OpenSourceJev25.869.150.9$0.011*Outside: latency 2.3× Jev
  33. –Raw Qwen3 8B direct logits51.145.550.2$0.065*Outside: cost 2.0× Jev
  34. –system-one50.545.049.9$0.068*Outside: cost 2.1× Jev
  35. –jeff4.481.142.4$0.0043*Outside: latency 11.6× Jev
  36. –Laya typed-decisions0.082.541.7$0.0038*Outside: latency 5.4× Jev
  37. –Bosun v3.1 0.6B13.677.539.4$0.0056*Outside: latency 6.5× Jev
  38. –open-jev-deberta-v3-large0.077.638.5$0.0056*Outside: latency 4.8× Jev
  39. –Laya0.084.936.9$0.0032*Outside: latency 2.4× Jev
  40. –Qwen3.5-0.8B Decision Model0.079.536.8$0.0048*Outside: latency 2.1× Jev
  41. –GLiNER2.5 multi14.086.636.0$0.0028*Outside: latency 2.2× Jev
  42. –GLiNER2 large22.577.632.4$0.0056*Outside: latency 3.0× Jev
  43. –SimpleJev16.770.331.7$0.0098*Outside: latency 12.0× Jev
  44. –Mirror5.689.324.4$0.0023*Outside: latency 7.4× Jev
  45. –Needle 30.061.60.0$0.019*Outside: latency 219.8× Jev
  46. –Needle 3, options as tools0.061.60.0$0.019*Outside: latency 94.3× Jev

Jev-class = cost per decision at most 2× Jev 1.13.0's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev 1.13.0's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 60 of 106 systems qualify; the other 46, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.

Capability against cost and speed

Jev-class systems are shown by default. Bubble size follows the official JevBench Score. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 60 systems. Upper right is best: more capable and cheaper.102030405060708090100$0.0010$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev← priciercheaper →1. Jev 1.13.02. Winnow-12B Q83. Cygnet4. Surogate Rune 26B-A4B v35. Jev-Omni
60 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 60 systems. Upper right is best: more capable and faster.10203040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev latency← slowerfaster →1. Jev 1.13.02. Winnow-12B Q83. Cygnet4. Surogate Rune 26B-A4B v35. Jev-Omni
60 systems. Tap a bubble for its values.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
  • faint = outside Jev-class

What-If: the weight sliders below re-score every system under other axis weights — only the equal 25/25/25/25 weights give the official option-A ranking. The 3D view of capability, cost and speed loads further down.

JevBench v1.5.4

JevBench Composite Score: 106 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

Weights:
Adjust weights ↓
View by:Capability ↑

Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).

Whiskers are 95% bootstrap intervals. 69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other 28 adjacent pairs, so no tie classification is inferred.

Greener = stronger within its column.

106 of 106 systems, sorted by official rank, #1 first.

  1. 1CygnetBase model: google/gemma-4-12B-itsource73.7I 71C 87S 91K 56est.$0.028
  2. 2Winnow-12B Q8Base model: google/gemma-4-12B-itsource73.2I 74C 84S 86K 57est.$0.028
  3. 3Jev 1.13.0APIBase model: undisclosed72.1I 72C 88S 84K 55est.$0.032
  4. 4JevK5 v0.3Base model: Qwen/Qwen3.5-4Bsource71.9I 56C 88S 94K 63est.$0.017
  5. 5Plumb-4BBase model: alibiserikbay/JevK5source71.6I 56C 87S 93K 63est.$0.017
  6. 6Jev-OmniBase model: google/gemma-4-12B-itsource71.5I 70C 83S 85K 56est.$0.029
  7. 7decider-4b v2Base model: Qwen/Qwen3.5-4B-Basesource71.3I 56C 86S 91K 65est.$0.015
  8. 8Decision 4B v1.2Base model: undisclosed70.8I 54C 89S 94K 63est.$0.017
  9. 9Imajev-4BBase model: Qwen/Qwen3.5-4Bsource70.4I 53C 88S 91K 63est.$0.017
  10. 10Decision 4B v1.1Base model: undisclosed70.4I 53C 87S 94K 63est.$0.017
  11. 11Manchego v2.1newBase model: Qwen/Qwen3.5-4Bsource68.8I 51C 85S 89K 64est.$0.015
  12. 12SemIfBase model: Qwen/Qwen3.5-4Bsource68.7I 51C 84S 91K 63est.$0.017
  13. 13spark-s1-4b-v6Base model: Qwen3.5-4Bsource68.2I 62C 70S 86K 60est.$0.021
  14. 14metask-jev-4bBase model: Qwen3.5-4Bsource67.5I 53C 83S 89K 58est.$0.026
  15. 15HopperBase model: Qwen/Qwen3.5-4Bsource67.5I 50C 88S 87K 62est.$0.018
  16. 16Malkuth-4BBase model: Qwen/Qwen3.5-4B-Basesource66.8I 54C 83S 86K 56est.$0.030
  17. 17Surogate Rune 26B-A4B v3Base model: google/gemma-4-26B-A4B-itsource66.5I 70C 88S 86K 49est.$0.050
  18. 18reflex 4BBase model: Qwen/Qwen3.5-4Bsource65.2I 51C 87S 69K 63est.$0.017
  19. 19jev-localBase model: Qwen/Qwen3.5-9Bsource65.2I 56C 78S 73K 59est.$0.024
  20. 20djev (Maisa, diffusion-gemma)Base model: google/diffusiongemma-26B-A4B-itsource64.2I 72C 80S 91K 48est.$0.053
Show all 106 systems (86 more ranked, 0 more not ranked)
  1. 21Raw Qwen3 4B Instruct 2507 direct logitsBase model: undisclosedsource62.1I 54C 53S 89K 63est.$0.017
  2. 22jqvBase model: Qwen3-32Bsource60.8I 49C 87S 83K 51est.$0.042
  3. 23JevK5 v0.2.0Base model: Qwen3.5-4Bsource58.1I 47C 85S 91K 63est.$0.017
  4. 24Qwen3.5-9B Jev-like data-mix v2Base model: Qwen/Qwen3.5-9Bsource53.0I 60C 81S 82K 46est.$0.065
  5. 25Standard One 8BBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source47.8I 60C 83S 92K 43est.$0.078
  6. 26NInfer Qwen3.8-Flash-Next mixedBase model: Qwen3.8-Flash-Nextsource47.5I 67C 89S 89K 43est.$0.082
  7. 27Instinct Dual 4BAPIBase model: Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.47.0I 43C 88S 82K 60est.$0.021
  8. 28swanOneBase model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source46.6I 71C 87S 85K 42est.$0.085
  9. 29Raw Qwen3 8B direct logitsBase model: Qwen/Qwen3-8B-Basesource45.2I 51C 49S 87K 46est.$0.065
  10. 30decider-2bBase model: Qwen/Qwen3.5-2B-Basesource45.1I 42C 72S 94K 65est.$0.015
  11. 31system-oneBase model: Qwen3-8Bsource44.1I 50C 49S 91K 45est.$0.068
  12. 32system-one-openAPIBase model: Gemma 4 E2Bsource42.4I 42C 72S 78K 68est.$0.011
  13. 33Autoloops – Gemma 4 31B ITAPIBase model: undisclosed40.5I 77C 86S 84K 40$0.103
  14. 34GPT-6 Luna (low reasoning effort)APIBase model: undisclosed40.5I 95C 95S 73K 39est.$0.108
  15. 35GPT-6 Luna (default medium reasoning effort)APIBase model: undisclosed38.8I 96C 96S 73K 38est.$0.114
  16. 36JevOneBase model: Qwen/Qwen3.6-35B-A3Bsource38.2I 54C 85S 90K 40est.$0.101
  17. 37kev 4BBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.38.1I 40C 68S 85K 66est.$0.014
  18. 38kev 8BBase model: Qwen/Qwen3-8B-Basesource34.2I 48C 71S 84K 40est.$0.097
  19. 39open-alternative-jevBase model: Qwen3.5-4Bsource33.6I 37C 77S 91K 63est.$0.017
  20. 40Bespoke Nimble 9BBase model: Qwen3.5-9Bsource31.8I 64C 77S 83K 37est.$0.128
  21. 41Malkuth-2BBase model: empero-ai/Qwen3.8-2B-Distillsource29.9I 36C 75S 92K 66est.$0.014
  22. 42openjev-sglangAPIBase model: Qwen/Qwen3.6-35B-A3Bsource29.0I 59C 83S 78K 36est.$0.140
  23. 43decider-35b-a3bBase model: Qwen/Qwen3.5-35B-A3B-Basesource27.5I 60C 82S 91K 34est.$0.154
  24. 44local-jev Qwen3.5-4BBase model: Qwen/Qwen3.5-4Bsource25.8I 34C 82S 84K 59est.$0.023
  25. 45Nemotron Diffusion 8BBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource25.7I 34C 76S 94K 55est.$0.030
  26. 46Open-Jev 9BBase model: Qwen/Qwen3.5-9Bsource24.4I 64C 82S 74K 33est.$0.170
  27. 47Decision 2BBase model: openbmb/MiniCPM5-2Bsource22.5I 31C 86S 90K 66est.$0.013
  28. 48GPT-5.6 LunaAPIBase model: undisclosed22.4I 94C 95S 74K 31$0.205
  29. 49typecastlmBase model: Qwen3.5-4Bsource21.8I 31C 77S 92K 64est.$0.016
  30. 50JEV Qwen3.5-9B Base NVFP4Base model: ig1/Qwen3.5-9B-NVFP4source20.1I 32C 81S 94K 48est.$0.056
  31. 51Gemini 3.1 Flash-LiteAPIBase model: undisclosed19.6I 78C 75S 80K 30$0.219
  32. 52AutoJev-27B (denis-pplx, Qwen3.8-27B)Base model: Qwen/Qwen3.8-27Bsource19.5I 73C 88S 88K 29est.$0.226
  33. 53AutoJev-27B (RTX PRO 6000)Base model: Qwen/Qwen3.8-27Bsource19.5I 73C 87S 87K 29est.$0.226
  34. 54NInfer Qwen3.8-27B NVFP4Base model: Qwen3.8-27Bsource18.7I 65C 86S 90K 29est.$0.231
  35. 55Eikos-27BBase model: Qwen/Qwen3.8-27Bsource18.5I 75C 86S 88K 29est.$0.238
  36. 56NInfer Qwen3.8-27B NVFP4 (T=1.5)Base model: Qwen3.8-27Bsource18.5I 61C 86S 90K 29est.$0.231
  37. 57InstinctAPIBase model: Qwen3.8-27Bsource · Operator-reported; weights not publicly verifiable.18.3I 63C 85S 82K 29est.$0.230
  38. 58OpenJevBase model: undisclosedsource17.9I 84C 83S 74K 29est.$0.241
  39. 59djev (thinking)Base model: google/diffusiongemma-26B-A4B-itsource17.4I 77C 96S 72K 28est.$0.249
  40. 60LitJevBase model: Qwen/Qwen3.8-27Bsource16.3I 58C 84S 68K 28est.$0.244
  41. 61Bev / Bonsai 27BBase model: Qwen/Qwen3.8-27Bsource15.8I 53C 78S 73K 28est.$0.247
  42. 62Raw Phi-4 mini direct logitsBase model: undisclosedsource15.2I 28C 71S 89K 53est.$0.036
  43. 63OpenSourceJevBase model: Qwen3.5-4Bsource13.2I 26C 76S 73K 69est.$0.011
  44. 64reflex-27bBase model: Qwen3.8-27Bsource13.2I 63C 86S 69K 26est.$0.297
  45. 65Open-Jev 2BBase model: Qwen/Qwen3.5-2Bsource9.1I 34C 74S 76K 33est.$0.170
  46. 66GLiNER2 largeBase model: microsoft/deberta-v3-largesource8.4I 22C 42S 65K 78est.$0.0056
  47. 67Qwen3-Reranker-4BBase model: Qwen/Qwen3-4B-Basesource7.1I 21C 76S 80K 48est.$0.052
  48. 68DeepSeek V4.1 FlashAPIBase model: undisclosed6.6I 94C 97S 69K 19est.$0.498
  49. 69SimpleJevBase model: Qwen/Qwen3.5-0.8Bsource4.0I 17C 47S 59K 70est.$0.0098
  50. 70SimpleJev Qwen3.8-27BAPIBase model: Qwen/Qwen3.8-27Bsource3.4I 73C 87S 75K 15est.$0.687
  51. 71decision-machine-1APIBase model: undisclosed3.2I 15C 81S 93K 56est.$0.029
  52. 72GLiNER2.5 multiBase model: mDeBERTa-v3-basesource2.7I 14C 58S 67K 87est.$0.0028
  53. 73Bosun v3.1 0.6BBase model: Qwen/Qwen3-0.6Bsource2.5I 14C 65S 62K 78est.$0.0056
  54. 74GLiNER2Base model: DeBERTa-v3-basesource2.3I 13C 36S 70K 87est.$0.0028
  55. 75Deem 0.8B v1Base model: Qwen/Qwen3.5-0.8Bsource2.1I 13C 39S 84K 80est.$0.0045
  56. 76JevActAPIBase model: undisclosedsource1.5I 11C 63S 76K 69est.$0.011
  57. 77CLM-8BBase model: Qwen/Qwen3-8Bsource1.5I 11C 48S 93K 51est.$0.045
  58. 78kev 0.6BBase model: Qwen/Qwen3-0.6B-Basesource1.3I 10C 68S 87K 80est.$0.0046
  59. 79Raw Qwen3 0.6B direct logitsBase model: Qwen/Qwen3-0.6B-Basesource1.1I 11C 21S 90K 78est.$0.0056
  60. 80GLiNER2.5 smallBase model: microsoft/deberta-v3-xsmallsource0.9I 9C 56S 77K 87est.$0.0028
  61. 81Raw Qwen3 1.7B direct logitsBase model: Qwen/Qwen3-1.7B-Basesource0.9I 10C 22S 90K 69est.$0.011
  62. 82MirrorBase model: undisclosed0.2I 6C 43S 64K 89est.$0.0023
  63. 83ZeroEntropy zerank-2Base model: Qwen/Qwen3-4Bsource0.1I 5C 82S 80K 48est.$0.052
  64. 84jeffBase model: undisclosedsource0.1I 4C 80S 56K 81est.$0.0043
  65. 85smalljev semantic-v9Base model: openbmb/MiniCPM5-2B-Basesource0.1I 4C 73S 86K 61est.$0.020
  66. 86Laya multilingualBase model: mmBERT-basesource0.0I 2C 44S 74K 82est.$0.0039
  67. 87OpenDecisionBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.0I 2C 73S 87K 79est.$0.0050
  68. 88BAAI bge-reranker-v2-m3Base model: BAAI/bge-m3source0.0I 0C 83S 91K 59est.$0.023
  69. 89Certo v1Base model: ModernBERT-largesource0.0I 0C 88S 91K 97est.$0.0013
  70. 90Decision FastBase model: Qwen/Qwen3-0.6B-Basesource0.0I 0C 76S 91K 80est.$0.0046
  71. 91Alibaba GTE Reranker ModernBERT-baseBase model: answerdotai/ModernBERT-basesource0.0I 0C 75S 91K 69est.$0.011
  72. 92kev 0.5BBase model: Qwen/Qwen2.5-0.5Bsource0.0I 0C 65S 88K 80est.$0.0046
  73. 93LayaBase model: ModernBERT-largesource0.0I 0C 74S 74K 85est.$0.0032
  74. 94lev-350mBase model: LiquidAI/LFM2.5-350Msource0.0I 0C 78S 94K 80est.$0.0046
  75. 95Qwen3.5-0.8B Decision ModelBase model: Qwen/Qwen3.5-0.8B-Basesource0.0I 0C 74S 72K 80est.$0.0048
  76. 96Mixedbread mxbai-rerank-base-v2Base model: undisclosedsource0.0I 0C 87S 89K 60est.$0.021
  77. 97Needle 3Base model: undisclosedsource0.0I 0C 0S 34K 62est.$0.019
  78. 98Needle 3, options as toolsBase model: undisclosedsource0.0I 0C 0S 41K 62est.$0.019
  79. 99open-jev-deberta-v3-largeBase model: microsoft/deberta-v3-largesource0.0I 0C 77S 68K 78est.$0.0056
  80. 100Open Jev JSON CanvasBase model: google/diffusiongemma-26B-A4B-itsource0.0I 77C 0S 86K 49est.$0.049
  81. 101openJev VerdictBase model: undisclosedsource0.0I 0C 52S 84K 87est.$0.0028
  82. 102openJev Verdict 1.4Base model: knowledgator/gliclass-modern-base-v2.0source0.0I 0C 80S 81K 87est.$0.0028
  83. 103Qwen3.8 27BAPIBase model: undisclosed0.0I 96C 98S 57K 0est.$2.178
  84. 104verdict-smallBase model: intfloat/multilingual-e5-smallsource0.0I 0C 59S 82K 100est.$0.0009
  85. 105VonBase model: answerdotai/ModernBERT-largesource0.0I 0C 83S 76K 83est.$0.0038
  86. 106Laya typed-decisionsBase model: ModernBERT-largesource0.0I 0C 83S 63K 83est.$0.0038
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; new = first listed in v1.5.4; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

Honorable mentions and Jev wrappers — listed separately, not ranked (1)

Eligibility rule: services that run on Jev itself may be measured and shown as honorable mentions, but are not competitors ranked against Jev and do not enter the field median gap (G_med) or tie markers. classifier.dev (TypeSafe) runs on Jev, so it stays unranked under this rule.

  • classifier.dev (fast tier)APIhonorable mention: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2). Official (A) score 74.7.

Compare two systems

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 72.1 (#3)
  • B: Cygnet — system-one-open · Score 73.7 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Cygnet. Intelligence: 72.0 vs 71.1; Calibration: 88.0 vs 87.0; Speed: 83.8 vs 91.0; Cost: 54.7 vs 56.4.50100Intelligence72.0 · 71.1Calibration88.0 · 87.0Speed83.8 · 91.0Cost54.7 · 56.4
0–100, the values in the table. A system with no published axis draws at 0 and says so.

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Jev 1.13.0 vs Cygnet. Math & numbers: 40.5 vs 31.4; Coding & software: 80.3 vs 87.8; Rules, policy & law: 73.7 vs 76.2; Finance & commerce: 83.4 vs 78.5; Support & operations: 91.5 vs 88.2; Everyday language: 96.1 vs 95.6; Safety & security: 67.7 vs 74.4.50100Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 235 items (174 open, 61 sealed).Math &numbers40.5 · 31.4Coding & software: code, SQL, repositories, developer tools and IT systems. 92 items (67 open, 25 sealed).Coding &software80.3 · 87.8Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 548 items (260 open, 288 sealed).Rules, policy& law73.7 · 76.2Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 224 items (134 open, 90 sealed).Finance &commerce83.4 · 78.5Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 282 items (131 open, 151 sealed).Support &operations91.5 · 88.2Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 115 items (98 open, 17 sealed).Everydaylanguage96.1 · 95.6Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 128 items (40 open, 88 sealed).Safety &security67.7 · 74.4
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.
What each category means · items per category
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 235 items (174 open / 61 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 92 items (67 open / 25 sealed)
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 548 items (260 open / 288 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 224 items (134 open / 90 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 282 items (131 open / 151 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 115 items (98 open / 17 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 128 items (40 open / 88 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jev 1.13.0 vs Cygnet. Search & retrieval: 95.9 vs 94.7; Model routing: 96.5 vs 99.1; LLM guardrails: 67.0 vs 72.3; Lead generation: 88.1 vs 82.2; Customer support: 91.9 vs 88.9; Insurance claims: 67.5 vs 71.7; Financial crime: 58.3 vs 71.1; Legal & compliance: 74.0 vs 75.3; E-commerce: 89.8 vs 91.6; Risk assessment: 80.1 vs 73.1; Knowledge graphs: 100.0 vs 73.2.50100Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 54 items (18 open, 36 sealed).Search &retrieval95.9 · 94.7Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 132 items (131 open, 1 sealed).Model routing96.5 · 99.1LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 339 items (209 open, 130 sealed).LLMguardrails67.0 · 72.3Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 98 items (61 open, 37 sealed).Leadgeneration88.1 · 82.2Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 389 items (176 open, 213 sealed).Customersupport91.9 · 88.9Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 31 items (14 open, 17 sealed).Insuranceclaims67.5 · 71.7Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 31 items (10 open, 21 sealed).Financialcrime58.3 · 71.1Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 508 items (261 open, 247 sealed).Legal &compliance74.0 · 75.3E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 71 items (48 open, 23 sealed).E-commerce89.8 · 91.6Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 78 items (39 open, 39 sealed).Riskassessment80.1 · 73.1Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 15 items (11 open, 4 sealed).Knowledgegraphs100.0 · 73.2
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.
What each category means · items per category
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 54 items (18 open / 36 sealed)
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 132 items (131 open / 1 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 339 items (209 open / 130 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 98 items (61 open / 37 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 389 items (176 open / 213 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 31 items (14 open / 17 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 31 items (10 open / 21 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 508 items (261 open / 247 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 71 items (48 open / 23 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 78 items (39 open / 39 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 15 items (11 open / 4 sealed)

Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 159 items (76 open / 83 sealed) — not a use case of its own, so it is counted but not drawn.

Low n (under 15 items, not plotted): Scientific discovery (1), Semantic code linting (4), Feature extraction for predictive modeling (7), Recruiting (3), Moderation and trust and safety (12), Advertising (2), Gaming (0), Demand forecasting (0).

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs Cygnet. Choice · open: 85.7 vs 82.9; Choice · sealed: 87.6 vs 80.3; Noul · open: 47.8 vs 57.9; Noul · sealed: 48.6 vs 55.7; Score · open: 81.2 vs 76.6; Score · sealed: 81.1 vs 73.2.50100Choice · open85.7 · 82.9Choice ·sealed87.6 · 80.3Noul · open47.8 · 57.9Noul · sealed48.6 · 55.7Score · open81.2 · 76.6Score ·sealed81.1 · 73.2
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 904 open and 720 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs Cygnet. Easy: 90.0 vs 89.7; Standard: 80.9 vs 82.4; Judge: 83.5 vs 82.2; Hard: 53.8 vs 58.0.50100Easy90.0 · 89.7Standard80.9 · 82.4Judge83.5 · 82.2Hard53.8 · 58.0
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs Cygnet. Easy: 91.2 vs 88.6; Standard: 79.0 vs 73.5; Judge: 73.7 vs 75.4; Hard: 72.9 vs 65.3.50100Easy91.2 · 88.6Standard79.0 · 73.5Judge73.7 · 75.4Hard72.9 · 65.3
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items, open and sealed pooled, all tiers: per request type the method's competence cell (Choice and Noul accuracy above chance, Score 1 - error / chance error), then averaged over types weighted by item count. Not part of the JevBench Score; compare systems within a category, not categories with each other. Values can be negative (below chance).

Every one of the 1,624 decisions was labelled with one subject topic (the v1.2 topic list, unchanged) and one or two TypeSafe use-case categories (docs.typesafe.ai/concepts/use-case-map, plus Other) by Winnow-12B Q8 (the model behind System1 Models s1-pro) on our own GPU pod; 1,118 of the 3,248 calls were first made on s1-pro itself and agree 98.2 % with the pod run. Sealed items stayed on our own infrastructure. About 5 % of public items were checked by hand (Claude Opus 5.5); the item-group rules below fix the systematic misses found.

  • routing, routing_hard and tool_selection items (choose the model, agent or tool that handles a request) are Model routing first; the labeller tagged many by the request subject.
  • tool_guardrail, unsafe_action, adequacy, judge_hard and policy_compliance items (check an agent tool call or a drafted answer before it goes out) are LLM guardrails first.
  • lead_qualification items take the topic Finance & commerce (sales); the labeller read the scoring rubric as rules & law.
  • A second use case is kept when the labeller gave it probability >= 0.15 and it is not Other.
All values as a table
SpokeA: Jev 1.13.0B: Cygnet
The four score axes
Intelligence72.071.1
Calibration88.087.0
Speed83.891.0
Cost54.756.4
Capability by subject topic
Math & numbers40.531.4
Coding & software80.387.8
Rules, policy & law73.776.2
Finance & commerce83.478.5
Support & operations91.588.2
Everyday language96.195.6
Safety & security67.774.4
Use cases (TypeSafe categories)
Search & retrieval95.994.7
Model routing96.599.1
LLM guardrails67.072.3
Lead generation88.182.2
Customer support91.988.9
Insurance claims67.571.7
Financial crime58.371.1
Legal & compliance74.075.3
E-commerce89.891.6
Risk assessment80.173.1
Knowledge graphs100.073.2
Competence per request type, open / sealed
Choice · open85.782.9
Choice · sealed87.680.3
Noul · open47.857.9
Noul · sealed48.655.7
Score · open81.276.6
Score · sealed81.173.2
Competence per tier — open set
Easy90.089.7
Standard80.982.4
Judge83.582.2
Hard53.858.0
Competence per tier — sealed set
Easy91.288.6
Standard79.073.5
Judge73.775.4
Hard72.965.3

Axes, request types, latency and cost

Compare:

Every measured system, under the four score axes. Intelligence is 50% open (904 decisions) and 50% sealed (720); Gap = I_open − I_sealed, and the penalty applies only above the field median gap (G_med 5.2) plus 8. The per-type competence, latency and price of the same systems are in Types & cost. On a phone the name column stays put while the table scrolls sideways.

#ASystemScoreIntel.Calib.SpeedCostGapPenalty
1CygnetdetailsBase model: google/gemma-4-12B-itsource73.771.187.091.056.4+2.7×1.000
2Winnow-12B Q8detailsBase model: google/gemma-4-12B-itsource73.274.484.186.156.6-3.9×1.000
3Jev 1.13.0APIdetailsBase model: undisclosed72.172.088.083.854.7-0.9×1.000
4JevK5 v0.3v1.5 roster addendum A1Base model: Qwen/Qwen3.5-4Bsource71.956.388.393.663.1+10.7×1.000
5Plumb-4Bv1.5 roster addendum A1Base model: alibiserikbay/JevK5source71.655.887.493.563.1+10.1×1.000
6Jev-OmnidetailsBase model: google/gemma-4-12B-itsource71.570.582.684.756.1-0.0×1.000
7decider-4b v2detailsBase model: Qwen/Qwen3.5-4B-Basesource71.355.885.690.964.5+11.9×1.000
8Decision 4B v1.2v1.5 roster addendum A1Base model: undisclosed70.853.788.693.563.1+5.2×1.000
9Imajev-4Bv1.5 roster addendum A2Base model: Qwen/Qwen3.5-4Bsource70.453.588.191.163.3+5.5×1.000
10Decision 4B v1.1v1.5 roster addendum A1Base model: undisclosed70.453.187.393.563.1+3.1×1.000
11Manchego v2.1v1.5 roster addendum A5detailsBase model: Qwen/Qwen3.5-4Bsource68.851.284.988.664.3+6.4×1.000
12SemIfdetailsBase model: Qwen/Qwen3.5-4Bsource68.751.384.090.963.1+5.4×1.000
13spark-s1-4b-v6detailsBase model: Qwen3.5-4Bsource68.262.169.785.560.4+0.7×1.000
14metask-jev-4bdetailsBase model: Qwen3.5-4Bsource67.553.582.789.557.7+6.5×1.000
15HopperdetailsBase model: Qwen/Qwen3.5-4Bsource67.549.987.987.262.3-2.3×1.000
16Malkuth-4BdetailsBase model: Qwen/Qwen3.5-4B-Basesource66.854.583.286.255.8+7.8×1.000
17Surogate Rune 26B-A4B v3v1.5 roster addendum A2Base model: google/gemma-4-26B-A4B-itsource66.569.788.386.049.0-1.6×1.000
18reflex 4BdetailsBase model: Qwen/Qwen3.5-4Bsource65.251.586.868.863.1+7.4×1.000
19jev-localdetailsBase model: Qwen/Qwen3.5-9Bsource65.256.377.873.058.9+3.8×1.000
20djev (Maisa, diffusion-gemma)detailsBase model: google/diffusiongemma-26B-A4B-itsource64.272.380.491.048.2+0.8×1.000
21Raw Qwen3 4B Instruct 2507 direct logitsdetailsBase model: undisclosedsource62.154.152.989.263.4+6.7×1.000
22jqvdetailsBase model: Qwen3-32Bsource60.849.186.883.451.2+4.5×1.000
23JevK5 v0.2.0detailsBase model: Qwen3.5-4Bsource58.146.784.990.963.1+6.6×1.000
24Qwen3.5-9B Jev-like data-mix v2detailsBase model: Qwen/Qwen3.5-9Bsource53.060.480.982.045.7+4.3×1.000
25Standard One 8BdetailsBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source47.859.683.292.543.3+0.7×1.000
26NInfer Qwen3.8-Flash-Next mixeddetailsBase model: Qwen3.8-Flash-Nextsource47.567.288.588.642.5+6.1×1.000
27Instinct Dual 4Bv1.5 roster addendum A4APIdetailsBase model: Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.47.043.188.382.060.2+6.3×1.000
28swanOnedetailsBase model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source46.671.287.184.742.2-1.5×1.000
29Raw Qwen3 8B direct logitsdetailsBase model: Qwen/Qwen3-8B-Basesource45.251.149.286.845.5+11.3×1.000
30decider-2bdetailsBase model: Qwen/Qwen3.5-2B-Basesource45.142.371.594.464.9+12.9×1.000
31system-onedetailsBase model: Qwen3-8Bsource44.150.549.490.945.0+6.1×1.000
32system-one-openAPIdetailsBase model: Gemma 4 E2Bsource42.441.671.678.368.3+5.6×1.000
33Autoloops – Gemma 4 31B ITAPIdetailsBase model: undisclosed40.576.785.883.939.6+0.4×1.000
34GPT-6 Luna (low reasoning effort)APIdetailsBase model: undisclosed40.595.394.973.239.1-1.1×1.000
35GPT-6 Luna (default medium reasoning effort)APIdetailsBase model: undisclosed38.896.295.673.238.3+0.2×1.000
36JevOnedetailsBase model: Qwen/Qwen3.6-35B-A3Bsource38.254.184.990.439.8-0.5×1.000
37kev 4BdetailsBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.38.139.967.685.565.8+13.9×0.993
38kev 8BdetailsBase model: Qwen/Qwen3-8B-Basesource34.248.371.084.140.4+12.0×1.000
39open-alternative-jevdetailsBase model: Qwen3.5-4Bsource33.637.476.991.363.3+10.5×1.000
40Bespoke Nimble 9BdetailsBase model: Qwen3.5-9Bsource31.863.777.283.036.8+2.8×1.000
41Malkuth-2BdetailsBase model: empero-ai/Qwen3.8-2B-Distillsource29.935.575.291.865.6+15.6×0.976
42openjev-sglangAPIdetailsBase model: Qwen/Qwen3.6-35B-A3Bsource29.058.682.978.135.6+5.4×1.000
43decider-35b-a3bdetailsBase model: Qwen/Qwen3.5-35B-A3B-Basesource27.560.582.091.034.4+7.4×1.000
44local-jev Qwen3.5-4BdetailsBase model: Qwen/Qwen3.5-4Bsource25.833.782.483.759.3+5.9×1.000
45Nemotron Diffusion 8Bv1.5 roster addendum A3detailsBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource25.733.876.293.755.5+11.4×1.000
46Open-Jev 9BdetailsBase model: Qwen/Qwen3.5-9Bsource24.463.881.573.633.1-4.1×1.000
47Decision 2BdetailsBase model: openbmb/MiniCPM5-2Bsource22.531.386.290.466.4+6.7×1.000
48GPT-5.6 LunaAPIdetailsBase model: undisclosed22.494.394.774.130.7-2.3×1.000
49typecastlmdetailsBase model: Qwen3.5-4Bsource21.831.276.792.064.0-15.1×1.000
50JEV Qwen3.5-9B Base NVFP4detailsBase model: ig1/Qwen3.5-9B-NVFP4source20.132.280.993.547.6+1.3×1.000
51Gemini 3.1 Flash-LiteAPIdetailsBase model: undisclosed19.677.674.780.129.8-1.9×1.000
52AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource19.572.887.787.629.4-1.7×1.000
53AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2Base model: Qwen/Qwen3.8-27Bsource19.572.886.787.129.4-1.5×1.000
54NInfer Qwen3.8-27B NVFP4detailsBase model: Qwen3.8-27Bsource18.765.585.989.929.1+3.0×1.000
55Eikos-27Bv1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource18.575.186.387.628.7-1.9×1.000
56NInfer Qwen3.8-27B NVFP4 (T=1.5)detailsBase model: Qwen3.8-27Bsource18.561.186.589.929.1+2.7×1.000
57InstinctAPIdetailsBase model: Qwen3.8-27Bsource · Operator-reported; weights not publicly verifiable.18.362.785.081.729.2-1.7×1.000
58OpenJev (thinking, BF16)detailsBase model: undisclosedsource17.984.283.173.928.5-3.4×1.000
59djev (thinking)detailsBase model: google/diffusiongemma-26B-A4B-itsource17.477.395.772.328.1-1.6×1.000
60LitJevdetailsBase model: Qwen/Qwen3.8-27Bsource16.358.384.568.228.4-0.5×1.000
61Bev / Bonsai 27Bv1.5 roster addendum A4detailsBase model: Qwen/Qwen3.8-27Bsource15.853.177.772.728.2-8.8×1.000
62Raw Phi-4 mini direct logitsdetailsBase model: undisclosedsource15.227.671.489.253.2+18.3×0.949
63OpenSourceJevdetailsBase model: Qwen3.5-4Bsource13.225.876.072.569.1+4.2×1.000
64reflex-27bdetailsBase model: Qwen3.8-27Bsource13.262.885.969.225.8+0.6×1.000
65Open-Jev 2BdetailsBase model: Qwen/Qwen3.5-2Bsource9.133.673.675.833.1+10.4×1.000
66GLiNER2 largedetailsBase model: microsoft/deberta-v3-largesource8.422.542.464.877.6+7.1×1.000
67Qwen3-Reranker-4BdetailsBase model: Qwen/Qwen3-4B-Basesource7.121.076.379.648.4+17.0×0.962
68DeepSeek V4.1 FlashAPIdetailsBase model: undisclosed6.693.796.969.419.1-3.7×1.000
69SimpleJevdetailsBase model: Qwen/Qwen3.5-0.8Bsource4.016.746.759.170.3-6.8×1.000
70SimpleJev Qwen3.8-27BAPIdetailsBase model: Qwen/Qwen3.8-27Bsource3.472.887.274.914.9-1.7×1.000
71decision-machine-1APIdetailsBase model: undisclosed3.214.881.292.856.3+22.4×0.908
72GLiNER2.5 multidetailsBase model: mDeBERTa-v3-basesource2.714.058.066.686.6+3.6×1.000
73Bosun v3.1 0.6Bv1.5 roster addendum A4detailsBase model: Qwen/Qwen3-0.6Bsource2.513.665.261.877.5+17.8×0.954
74GLiNER2detailsBase model: DeBERTa-v3-basesource2.313.535.670.386.6+7.5×1.000
75Deem 0.8B v1v1.5 roster addendum A4detailsBase model: Qwen/Qwen3.5-0.8Bsource2.113.038.984.380.3+21.1×0.921
76JevActAPIdetailsBase model: undisclosedsource1.511.262.576.468.6+16.3×0.969
77CLM-8BdetailsBase model: Qwen/Qwen3-8Bsource1.511.448.593.350.5+11.3×1.000
78kev 0.6BdetailsBase model: Qwen/Qwen3-0.6B-Basesource1.310.367.787.280.1+27.0×0.862
79Raw Qwen3 0.6B direct logitsdetailsBase model: Qwen/Qwen3-0.6B-Basesource1.110.621.590.477.6-1.1×1.000
80GLiNER2.5 smalldetailsBase model: microsoft/deberta-v3-xsmallsource0.99.255.977.486.6+15.5×0.977
81Raw Qwen3 1.7B direct logitsdetailsBase model: Qwen/Qwen3-1.7B-Basesource0.99.721.690.268.5+5.2×1.000
82MirrordetailsBase model: undisclosed0.25.643.263.689.3+2.3×1.000
83ZeroEntropy zerank-2detailsBase model: Qwen/Qwen3-4Bsource0.14.881.780.348.4+10.4×1.000
84jeffdetailsBase model: undisclosedsource0.14.480.355.881.1+16.3×0.969
85smalljev semantic-v9detailsBase model: openbmb/MiniCPM5-2B-Basesource0.13.773.186.560.8+14.5×0.987
86Laya multilingualv1.5 roster addendum A4detailsBase model: mmBERT-basesource0.02.443.673.782.2+0.9×1.000
87OpenDecisiondetailsBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.02.172.586.679.1+14.7×0.985
88BAAI bge-reranker-v2-m3detailsBase model: BAAI/bge-m3source0.00.083.390.559.3+1.7×1.000
89Certo v1detailsBase model: ModernBERT-largesource0.00.087.591.496.9+0.8×1.000
90Decision FastdetailsBase model: Qwen/Qwen3-0.6B-Basesource0.00.076.191.480.1+19.0×0.942
91Alibaba GTE Reranker ModernBERT-basedetailsBase model: answerdotai/ModernBERT-basesource0.00.075.291.569.3+4.5×1.000
92kev 0.5BdetailsBase model: Qwen/Qwen2.5-0.5Bsource0.00.065.088.580.1+25.7×0.875
93LayadetailsBase model: ModernBERT-largesource0.00.073.773.984.9+12.7×1.000
94lev-350mdetailsBase model: LiquidAI/LFM2.5-350Msource0.00.077.793.680.0+23.5×0.896
95Qwen3.5-0.8B Decision ModeldetailsBase model: Qwen/Qwen3.5-0.8B-Basesource0.00.073.671.879.5+7.1×1.000
96Mixedbread mxbai-rerank-base-v2detailsBase model: undisclosedsource0.00.086.689.360.3+3.0×1.000
97Needle 3detailsBase model: undisclosedsource0.00.00.034.161.6+10.0×1.000
98Needle 3, options as toolsdetailsBase model: undisclosedsource0.00.00.041.061.6+14.1×0.991
99open-jev-deberta-v3-largedetailsBase model: microsoft/deberta-v3-largesource0.00.077.168.377.6+5.3×1.000
100Open Jev JSON CanvasdetailsBase model: google/diffusiongemma-26B-A4B-itsource0.077.10.085.649.3+4.0×1.000
101openJev VerdictdetailsBase model: undisclosedsource0.00.052.283.986.6+26.7×0.865
102openJev Verdict 1.4detailsBase model: knowledgator/gliclass-modern-base-v2.0source0.00.080.380.686.6+3.5×1.000
103Qwen3.8 27BAPIdetailsBase model: undisclosed0.095.698.156.80.0-5.3×1.000
104verdict-smalldetailsBase model: intfloat/multilingual-e5-smallsource0.00.059.582.1100.0+2.0×1.000
105VondetailsBase model: answerdotai/ModernBERT-largesource0.00.083.575.782.7+14.7×0.985
106Laya typed-decisionsv1.5 roster addendum A4detailsBase model: ModernBERT-largesource0.00.083.363.482.5+8.5×1.000
–classifier.devAPIhonorable mentiondetailsBase model: undisclosed—75.889.282.058.9-2.7×1.000
–SimpleJev Qwen3.6-35B-A3BAPIpartial rundetailsBase model: undisclosedsource—0.078.575.235.1+5.6×1.000
–Decision-4Bv1.5 roster addendum A4not rankeddetailsBase model: Qwen/Qwen3.5-4Bsource———————
Official order with 95% intervals (106 systems)

JevBench v1.5.4 · headline option A

JevBench Score: 106 ranked systems

Official (A)weighted harmonic mean of four 0–100 axes, Intelligence · Calibration · Speed · Cost = 25 · 25 · 25 · 25, with the low-axis gates · Method ↓

Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).

Whiskers are 95% bootstrap intervals. 69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other 28 adjacent pairs, so no tie classification is inferred.

  1. 1CygnetBase model: google/gemma-4-12B-itsource73.7I 71 · C 87 · S 91 · K 56 · B#2 · C#1 · $0.028
  2. 2Winnow-12B Q8Base model: google/gemma-4-12B-itsource73.2I 74 · C 84 · S 86 · K 57 · B#1 · C#2 · $0.028
  3. 3Jev 1.13.0APIBase model: undisclosed72.1I 72 · C 88 · S 84 · K 55 · B#3 · C#3 · $0.032
  4. 4JevK5 v0.3v1.5 roster addendum A1Base model: Qwen/Qwen3.5-4Bsource71.9I 56 · C 88 · S 94 · K 63 · B#5 · C#8 · $0.017
  5. 5Plumb-4Bv1.5 roster addendum A1Base model: alibiserikbay/JevK5source71.6I 56 · C 87 · S 93 · K 63 · B#6 · C#9 · $0.017
  6. 6Jev-OmniBase model: google/gemma-4-12B-itsource71.5I 70 · C 83 · S 85 · K 56 · B#4 · C#4 · $0.029
  7. 7decider-4b v2Base model: Qwen/Qwen3.5-4B-Basesource71.3I 56 · C 86 · S 91 · K 65 · B#7 · C#10 · $0.015
  8. 8Decision 4B v1.2v1.5 roster addendum A1Base model: undisclosed70.8I 54 · C 89 · S 94 · K 63 · B#9 · C#12 · $0.017
  9. 9Imajev-4Bv1.5 roster addendum A2Base model: Qwen/Qwen3.5-4Bsource70.4I 53 · C 88 · S 91 · K 63 · B#11 · C#13 · $0.017
  10. 10Decision 4B v1.1v1.5 roster addendum A1Base model: undisclosed70.4I 53 · C 87 · S 94 · K 63 · B#12 · C#14 · $0.017
  11. 11Manchego v2.1v1.5 roster addendum A5Base model: Qwen/Qwen3.5-4Bsource68.8I 51 · C 85 · S 89 · K 64 · B#14 · C#20 · $0.015
  12. 12SemIfBase model: Qwen/Qwen3.5-4Bsource68.7I 51 · C 84 · S 91 · K 63 · B#15 · C#19 · $0.017
  13. 13spark-s1-4b-v6Base model: Qwen3.5-4Bsource68.2I 62 · C 70 · S 86 · K 60 · B#8 · C#5 · $0.021
  14. 14metask-jev-4bBase model: Qwen3.5-4Bsource67.5I 53 · C 83 · S 89 · K 58 · B#16 · C#16 · $0.026
  15. 15HopperBase model: Qwen/Qwen3.5-4Bsource67.5I 50 · C 88 · S 87 · K 62 · B#19 · C#24 · $0.018
  16. 16Malkuth-4BBase model: Qwen/Qwen3.5-4B-Basesource66.8I 54 · C 83 · S 86 · K 56 · B#17 · C#15 · $0.030
  17. 17Surogate Rune 26B-A4B v3v1.5 roster addendum A2Base model: google/gemma-4-26B-A4B-itsource66.5I 70 · C 88 · S 86 · K 49 · B#10 · C#6 · $0.050
  18. 18reflex 4BBase model: Qwen/Qwen3.5-4Bsource65.2I 51 · C 87 · S 69 · K 63 · B#20 · C#21 · $0.017
  19. 19jev-localBase model: Qwen/Qwen3.5-9Bsource65.2I 56 · C 78 · S 73 · K 59 · B#18 · C#11 · $0.024
  20. 20djev (Maisa, diffusion-gemma)Base model: google/diffusiongemma-26B-A4B-itsource64.2I 72 · C 80 · S 91 · K 48 · B#13 · C#7 · $0.053
  21. 21Raw Qwen3 4B Instruct 2507 direct logitsBase model: undisclosedsource62.1I 54 · C 53 · S 89 · K 63 · B#21 · C#18 · $0.017
  22. 22jqvBase model: Qwen3-32Bsource60.8I 49 · C 87 · S 83 · K 51 · B#22 · C#26 · $0.042
  23. 23JevK5 v0.2.0Base model: Qwen3.5-4Bsource58.1I 47 · C 85 · S 91 · K 63 · B#23 · C#29 · $0.017
  24. 24Qwen3.5-9B Jev-like data-mix v2Base model: Qwen/Qwen3.5-9Bsource53.0I 60 · C 81 · S 82 · K 46 · B#24 · C#17 · $0.065
  25. 25Standard One 8BBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source47.8I 60 · C 83 · S 92 · K 43 · B#27 · C#23 · $0.078
  26. 26NInfer Qwen3.8-Flash-Next mixedBase model: Qwen3.8-Flash-Nextsource47.5I 67 · C 89 · S 89 · K 43 · B#25 · C#22 · $0.082
  27. 27Instinct Dual 4Bv1.5 roster addendum A4APIBase model: Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.47.0I 43 · C 88 · S 82 · K 60 · B#31 · C#32 · $0.021
  28. 28swanOneBase model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source46.6I 71 · C 87 · S 85 · K 42 · B#26 · C#25 · $0.085
  29. 29Raw Qwen3 8B direct logitsBase model: Qwen/Qwen3-8B-Basesource45.2I 51 · C 49 · S 87 · K 46 · B#28 · C#31 · $0.065
  30. 30decider-2bBase model: Qwen/Qwen3.5-2B-Basesource45.1I 42 · C 72 · S 94 · K 65 · B#34 · C#34 · $0.015
  31. 31system-oneBase model: Qwen3-8Bsource44.1I 50 · C 49 · S 91 · K 45 · B#29 · C#35 · $0.068
  32. 32system-one-openAPIBase model: Gemma 4 E2Bsource42.4I 42 · C 72 · S 78 · K 68 · B#35 · C#37 · $0.011
  33. 33Autoloops – Gemma 4 31B ITAPIBase model: undisclosed40.5I 77 · C 86 · S 84 · K 40 · B#32 · C#27 · tariff$0.103
  34. 34GPT-6 Luna (low reasoning effort)APIBase model: undisclosed40.5I 95 · C 95 · S 73 · K 39 · B#30 · C#28 · $0.108
  35. 35GPT-6 Luna (default medium reasoning effort)APIBase model: undisclosed38.8I 96 · C 96 · S 73 · K 38 · B#33 · C#30 · $0.114
  36. 36JevOneBase model: Qwen/Qwen3.6-35B-A3Bsource38.2I 54 · C 85 · S 90 · K 40 · B#36 · C#36 · $0.101
  37. 37kev 4BBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.38.1I 40 · C 68 · S 85 · K 66 · B#37 · C#40 · $0.014
  38. 38kev 8BBase model: Qwen/Qwen3-8B-Basesource34.2I 48 · C 71 · S 84 · K 40 · B#38 · C#42 · $0.097
  39. 39open-alternative-jevBase model: Qwen3.5-4Bsource33.6I 37 · C 77 · S 91 · K 63 · B#40 · C#43 · $0.017
  40. 40Bespoke Nimble 9BBase model: Qwen3.5-9Bsource31.8I 64 · C 77 · S 83 · K 37 · B#39 · C#33 · $0.128
  41. 41Malkuth-2BBase model: empero-ai/Qwen3.8-2B-Distillsource29.9I 36 · C 75 · S 92 · K 66 · B#43 · C#45 · $0.014
  42. 42openjev-sglangAPIBase model: Qwen/Qwen3.6-35B-A3Bsource29.0I 59 · C 83 · S 78 · K 36 · B#41 · C#38 · $0.140
  43. 43decider-35b-a3bBase model: Qwen/Qwen3.5-35B-A3B-Basesource27.5I 60 · C 82 · S 91 · K 34 · B#42 · C#39 · $0.154
  44. 44local-jev Qwen3.5-4BBase model: Qwen/Qwen3.5-4Bsource25.8I 34 · C 82 · S 84 · K 59 · B#46 · C#54 · $0.023
  45. 45Nemotron Diffusion 8Bv1.5 roster addendum A3Base model: nvidia/Nemotron-Labs-Diffusion-8Bsource25.7I 34 · C 76 · S 94 · K 55 · B#47 · C#55 · $0.030
  46. 46Open-Jev 9BBase model: Qwen/Qwen3.5-9Bsource24.4I 64 · C 82 · S 74 · K 33 · B#44 · C#41 · $0.170
  47. 47Decision 2BBase model: openbmb/MiniCPM5-2Bsource22.5I 31 · C 86 · S 90 · K 66 · B#53 · C#57 · $0.013
  48. 48GPT-5.6 LunaAPIBase model: undisclosed22.4I 94 · C 95 · S 74 · K 31 · B#45 · C#44 · tariff$0.205
  49. 49typecastlmBase model: Qwen3.5-4Bsource21.8I 31 · C 77 · S 92 · K 64 · B#57 · C#59 · $0.016
  50. 50JEV Qwen3.5-9B Base NVFP4Base model: ig1/Qwen3.5-9B-NVFP4source20.1I 32 · C 81 · S 94 · K 48 · B#59 · C#60 · $0.056
  51. 51Gemini 3.1 Flash-LiteAPIBase model: undisclosed19.6I 78 · C 75 · S 80 · K 30 · B#48 · C#46 · tariff$0.219
  52. 52AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource19.5I 73 · C 88 · S 88 · K 29 · B#49 · C#47 · $0.226
  53. 53AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2Base model: Qwen/Qwen3.8-27Bsource19.5I 73 · C 87 · S 87 · K 29 · B#50 · C#48 · $0.226
  54. 54NInfer Qwen3.8-27B NVFP4Base model: Qwen3.8-27Bsource18.7I 65 · C 86 · S 90 · K 29 · B#52 · C#49 · $0.231
  55. 55Eikos-27Bv1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource18.5I 75 · C 86 · S 88 · K 29 · B#51 · C#50 · $0.238
  56. 56NInfer Qwen3.8-27B NVFP4 (T=1.5)Base model: Qwen3.8-27Bsource18.5I 61 · C 86 · S 90 · K 29 · B#55 · C#51 · $0.231
  57. 57InstinctAPIBase model: Qwen3.8-27Bsource · Operator-reported; weights not publicly verifiable.18.3I 63 · C 85 · S 82 · K 29 · B#56 · C#52 · $0.230
  58. 58OpenJev (thinking, BF16)Base model: undisclosedsource17.9I 84 · C 83 · S 74 · K 29 · B#54 · C#53 · $0.241
  59. 59djev (thinking)Base model: google/diffusiongemma-26B-A4B-itsource17.4I 77 · C 96 · S 72 · K 28 · B#58 · C#56 · $0.249
  60. 60LitJevBase model: Qwen/Qwen3.8-27Bsource16.3I 58 · C 84 · S 68 · K 28 · B#60 · C#58 · $0.244
  61. 61Bev / Bonsai 27Bv1.5 roster addendum A4Base model: Qwen/Qwen3.8-27Bsource15.8I 53 · C 78 · S 73 · K 28 · B#61 · C#62 · $0.247
  62. 62Raw Phi-4 mini direct logitsBase model: undisclosedsource15.2I 28 · C 71 · S 89 · K 53 · B#63 · C#63 · $0.036
  63. 63OpenSourceJevBase model: Qwen3.5-4Bsource13.2I 26 · C 76 · S 73 · K 69 · B#64 · C#64 · $0.011
  64. 64reflex-27bBase model: Qwen3.8-27Bsource13.2I 63 · C 86 · S 69 · K 26 · B#62 · C#61 · $0.297
  65. 65Open-Jev 2BBase model: Qwen/Qwen3.5-2Bsource9.1I 34 · C 74 · S 76 · K 33 · B#65 · C#66 · $0.170
  66. 66GLiNER2 largeBase model: microsoft/deberta-v3-largesource8.4I 22 · C 42 · S 65 · K 78 · B#67 · C#67 · $0.0056
  67. 67Qwen3-Reranker-4BBase model: Qwen/Qwen3-4B-Basesource7.1I 21 · C 76 · S 80 · K 48 · B#68 · C#68 · $0.052
  68. 68DeepSeek V4.1 FlashAPIBase model: undisclosed6.6I 94 · C 97 · S 69 · K 19 · B#66 · C#65 · $0.498
  69. 69SimpleJevBase model: Qwen/Qwen3.5-0.8Bsource4.0I 17 · C 47 · S 59 · K 70 · B#70 · C#70 · $0.0098
  70. 70SimpleJev Qwen3.8-27BAPIBase model: Qwen/Qwen3.8-27Bsource3.4I 73 · C 87 · S 75 · K 15 · B#69 · C#69 · $0.687
  71. 71decision-machine-1APIBase model: undisclosed3.2I 15 · C 81 · S 93 · K 56 · B#71 · C#71 · $0.029
  72. 72GLiNER2.5 multiBase model: mDeBERTa-v3-basesource2.7I 14 · C 58 · S 67 · K 87 · B#72 · C#72 · $0.0028
  73. 73Bosun v3.1 0.6Bv1.5 roster addendum A4Base model: Qwen/Qwen3-0.6Bsource2.5I 14 · C 65 · S 62 · K 78 · B#73 · C#73 · $0.0056
  74. 74GLiNER2Base model: DeBERTa-v3-basesource2.3I 13 · C 36 · S 70 · K 87 · B#74 · C#74 · $0.0028
  75. 75Deem 0.8B v1v1.5 roster addendum A4Base model: Qwen/Qwen3.5-0.8Bsource2.1I 13 · C 39 · S 84 · K 80 · B#75 · C#75 · $0.0045
  76. 76JevActAPIBase model: undisclosedsource1.5I 11 · C 63 · S 76 · K 69 · B#77 · C#76 · $0.011
  77. 77CLM-8BBase model: Qwen/Qwen3-8Bsource1.5I 11 · C 48 · S 93 · K 51 · B#76 · C#77 · $0.045
  78. 78kev 0.6BBase model: Qwen/Qwen3-0.6B-Basesource1.3I 10 · C 68 · S 87 · K 80 · B#78 · C#78 · $0.0046
  79. 79Raw Qwen3 0.6B direct logitsBase model: Qwen/Qwen3-0.6B-Basesource1.1I 11 · C 21 · S 90 · K 78 · B#79 · C#79 · $0.0056
  80. 80GLiNER2.5 smallBase model: microsoft/deberta-v3-xsmallsource0.9I 9 · C 56 · S 77 · K 87 · B#81 · C#80 · $0.0028
  81. 81Raw Qwen3 1.7B direct logitsBase model: Qwen/Qwen3-1.7B-Basesource0.9I 10 · C 22 · S 90 · K 69 · B#80 · C#81 · $0.011
  82. 82MirrorBase model: undisclosed0.2I 6 · C 43 · S 64 · K 89 · B#82 · C#82 · $0.0023
  83. 83ZeroEntropy zerank-2Base model: Qwen/Qwen3-4Bsource0.1I 5 · C 82 · S 80 · K 48 · B#83 · C#83 · $0.052
  84. 84jeffBase model: undisclosedsource0.1I 4 · C 80 · S 56 · K 81 · B#84 · C#84 · $0.0043
  85. 85smalljev semantic-v9Base model: openbmb/MiniCPM5-2B-Basesource0.1I 4 · C 73 · S 86 · K 61 · B#85 · C#85 · $0.020
  86. 86Laya multilingualv1.5 roster addendum A4Base model: mmBERT-basesource0.0I 2 · C 44 · S 74 · K 82 · B#86 · C#86 · $0.0039
  87. 87OpenDecisionBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.0I 2 · C 73 · S 87 · K 79 · B#87 · C#87 · $0.0050
  88. 88BAAI bge-reranker-v2-m3Base model: BAAI/bge-m3source0.0I 0 · C 83 · S 91 · K 59 · B#88 · C#88 · $0.023
  89. 89Certo v1Base model: ModernBERT-largesource0.0I 0 · C 88 · S 91 · K 97 · B#89 · C#89 · $0.0013
  90. 90Decision FastBase model: Qwen/Qwen3-0.6B-Basesource0.0I 0 · C 76 · S 91 · K 80 · B#90 · C#90 · $0.0046
  91. 91Alibaba GTE Reranker ModernBERT-baseBase model: answerdotai/ModernBERT-basesource0.0I 0 · C 75 · S 91 · K 69 · B#91 · C#91 · $0.011
  92. 92kev 0.5BBase model: Qwen/Qwen2.5-0.5Bsource0.0I 0 · C 65 · S 88 · K 80 · B#92 · C#92 · $0.0046
  93. 93LayaBase model: ModernBERT-largesource0.0I 0 · C 74 · S 74 · K 85 · B#93 · C#93 · $0.0032
  94. 94lev-350mBase model: LiquidAI/LFM2.5-350Msource0.0I 0 · C 78 · S 94 · K 80 · B#94 · C#94 · $0.0046
  95. 95Qwen3.5-0.8B Decision ModelBase model: Qwen/Qwen3.5-0.8B-Basesource0.0I 0 · C 74 · S 72 · K 80 · B#95 · C#95 · $0.0048
  96. 96Mixedbread mxbai-rerank-base-v2Base model: undisclosedsource0.0I 0 · C 87 · S 89 · K 60 · B#96 · C#96 · $0.021
  97. 97Needle 3Base model: undisclosedsource0.0I 0 · C 0 · S 34 · K 62 · B#97 · C#97 · $0.019
  98. 98Needle 3, options as toolsBase model: undisclosedsource0.0I 0 · C 0 · S 41 · K 62 · B#98 · C#98 · $0.019
  99. 99open-jev-deberta-v3-largeBase model: microsoft/deberta-v3-largesource0.0I 0 · C 77 · S 68 · K 78 · B#99 · C#99 · $0.0056
  100. 100Open Jev JSON CanvasBase model: google/diffusiongemma-26B-A4B-itsource0.0I 77 · C 0 · S 86 · K 49 · B#100 · C#100 · $0.049
  101. 101openJev VerdictBase model: undisclosedsource0.0I 0 · C 52 · S 84 · K 87 · B#101 · C#101 · $0.0028
  102. 102openJev Verdict 1.4Base model: knowledgator/gliclass-modern-base-v2.0source0.0I 0 · C 80 · S 81 · K 87 · B#102 · C#102 · $0.0028
  103. 103Qwen3.8 27BAPIBase model: undisclosed0.0I 96 · C 98 · S 57 · K 0 · B#103 · C#103 · $2.18
  104. 104verdict-smallBase model: intfloat/multilingual-e5-smallsource0.0I 0 · C 59 · S 82 · K 100 · B#104 · C#104 · $0.0009
  105. 105VonBase model: answerdotai/ModernBERT-largesource0.0I 0 · C 83 · S 76 · K 83 · B#105 · C#105 · $0.0038
  106. 106Laya typed-decisionsv1.5 roster addendum A4Base model: ModernBERT-largesource0.0I 0 · C 83 · S 63 · K 83 · B#106 · C#106 · $0.0038

Costs are estimates (est.) unless marked tariff.

system-one-openJev rebuildJev (TypeSafe, closed)UnclassifiedRaw-logit control (base model)Native-logit decision engineClosed decision APIInstruction model, JSON schemaZero-shot classifierReranker (neutral adapter)Small tool-calling model

All three weight options (106 systems)

All three weight options

A is the official headline: equal 25/25/25/25 axis weights and an Intelligence floor of 50. B remains the secondary 40/20/20/20 axis-weight view; C retains equal axes with an Intelligence floor of 60. All three use equal Choice/Noul/Score weights. The CI column is the paired-bootstrap 95% interval of the A score.

#ASystemA · equal (headline)B · 40/20/20/20C · equal, I floor 60#B#CA 95% CI
1CygnetdetailsBase model: google/gemma-4-12B-itsource73.773.273.72172.4–74.5
2Winnow-12B Q8detailsBase model: google/gemma-4-12B-itsource73.273.573.21272.0–74.0
3Jev 1.13.0APIdetailsBase model: undisclosed72.172.172.13371.0–72.6
4JevK5 v0.3v1.5 roster addendum A1Base model: Qwen/Qwen3.5-4Bsource71.968.163.25869.4–72.9
5Plumb-4Bv1.5 roster addendum A1Base model: alibiserikbay/JevK5source71.667.762.06969.2–72.7
6Jev-OmnidetailsBase model: google/gemma-4-12B-itsource71.571.371.54470.2–72.4
7decider-4b v2detailsBase model: Qwen/Qwen3.5-4B-Basesource71.367.561.671069.1–72.3
8Decision 4B v1.2v1.5 roster addendum A1Base model: undisclosed70.866.656.791268.5–72.0
9Imajev-4Bv1.5 roster addendum A2Base model: Qwen/Qwen3.5-4Bsource70.466.255.9111367.8–71.6
10Decision 4B v1.1v1.5 roster addendum A1Base model: undisclosed70.466.155.1121466.9–71.6
11Manchego v2.1v1.5 roster addendum A5detailsBase model: Qwen/Qwen3.5-4Bsource68.864.450.1142059.6–70.3
12SemIfdetailsBase model: Qwen/Qwen3.5-4Bsource68.764.350.2151960.1–70.2
13spark-s1-4b-v6detailsBase model: Qwen3.5-4Bsource68.266.968.28566.3–69.7
14metask-jev-4bdetailsBase model: Qwen3.5-4Bsource67.564.153.6161665.4–68.7
15HopperdetailsBase model: Qwen/Qwen3.5-4Bsource67.562.946.9192456.3–69.2
16Malkuth-4BdetailsBase model: Qwen/Qwen3.5-4B-Basesource66.863.955.0171564.9–67.9
17Surogate Rune 26B-A4B v3v1.5 roster addendum A2Base model: google/gemma-4-26B-A4B-itsource66.566.666.510665.4–67.0
18reflex 4BdetailsBase model: Qwen/Qwen3.5-4Bsource65.261.948.0202158.6–66.3
19jev-localdetailsBase model: Qwen/Qwen3.5-9Bsource65.263.257.4181163.5–66.5
20djev (Maisa, diffusion-gemma)detailsBase model: google/diffusiongemma-26B-A4B-itsource64.264.864.213763.1–65.0
21Raw Qwen3 4B Instruct 2507 direct logitsdetailsBase model: undisclosedsource62.160.350.4211859.4–64.2
22jqvdetailsBase model: Qwen3-32Bsource60.857.542.2222651.3–64.3
23JevK5 v0.2.0detailsBase model: Qwen3.5-4Bsource58.153.640.4232947.6–68.2
24Qwen3.5-9B Jev-like data-mix v2detailsBase model: Qwen/Qwen3.5-9Bsource53.052.553.0241751.9–53.9
25Standard One 8BdetailsBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source47.847.147.2272346.7–48.5
26NInfer Qwen3.8-Flash-Next mixeddetailsBase model: Qwen3.8-Flash-Nextsource47.547.747.5252246.7–47.9
27Instinct Dual 4Bv1.5 roster addendum A4APIdetailsBase model: Qwen3.5-4Bsource · Operator-reported; weights not publicly verifiable.47.043.032.6313237.9–56.8
28swanOnedetailsBase model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source46.647.346.6262545.8–46.9
29Raw Qwen3 8B direct logitsdetailsBase model: Qwen/Qwen3-8B-Basesource45.244.632.9283136.9–46.7
30decider-2bdetailsBase model: Qwen/Qwen3.5-2B-Basesource45.141.131.3343433.8–54.3
31system-onedetailsBase model: Qwen3-8Bsource44.143.431.2293536.8–45.7
32system-one-openAPIdetailsBase model: Gemma 4 E2Bsource42.438.829.5353733.8–51.9
33Autoloops – Gemma 4 31B ITAPIdetailsBase model: undisclosed40.541.840.5322740.1–40.8
34GPT-6 Luna (low reasoning effort)APIdetailsBase model: undisclosed40.543.140.5302840.3–40.7
35GPT-6 Luna (default medium reasoning effort)APIdetailsBase model: undisclosed38.841.438.8333038.6–38.9
36JevOnedetailsBase model: Qwen/Qwen3.6-35B-A3Bsource38.237.431.1363637.4–38.8
37kev 4BdetailsBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.38.134.626.4374027.5–46.8
38kev 8BdetailsBase model: Qwen/Qwen3-8B-Basesource34.233.123.7384226.8–37.5
39open-alternative-jevdetailsBase model: Qwen3.5-4Bsource33.629.923.3404325.8–41.7
40Bespoke Nimble 9BdetailsBase model: Qwen3.5-9Bsource31.832.331.8393331.2–32.3
41Malkuth-2BdetailsBase model: empero-ai/Qwen3.8-2B-Distillsource29.926.420.8434521.0–39.0
42openjev-sglangAPIdetailsBase model: Qwen/Qwen3.6-35B-A3Bsource29.029.227.7413828.4–29.4
43decider-35b-a3bdetailsBase model: Qwen/Qwen3.5-35B-A3B-Basesource27.527.727.5423926.9–27.9
44local-jev Qwen3.5-4BdetailsBase model: Qwen/Qwen3.5-4Bsource25.822.717.9465419.9–33.0
45Nemotron Diffusion 8Bv1.5 roster addendum A3detailsBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource25.722.717.8475518.6–32.8
46Open-Jev 9BdetailsBase model: Qwen/Qwen3.5-9Bsource24.425.024.4444123.9–24.6
47Decision 2BdetailsBase model: openbmb/MiniCPM5-2Bsource22.519.315.6535716.7–28.8
48GPT-5.6 LunaAPIdetailsBase model: undisclosed22.424.222.4454422.3–22.5
49typecastlmdetailsBase model: Qwen3.5-4Bsource21.818.915.2575916.2–28.6
50JEV Qwen3.5-9B Base NVFP4detailsBase model: ig1/Qwen3.5-9B-NVFP4source20.117.813.9596015.2–25.5
51Gemini 3.1 Flash-LiteAPIdetailsBase model: undisclosed19.620.819.6484619.3–19.8
52AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource19.520.419.5494719.3–19.6
53AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2Base model: Qwen/Qwen3.8-27Bsource19.520.419.5504819.3–19.6
54NInfer Qwen3.8-27B NVFP4detailsBase model: Qwen3.8-27Bsource18.719.318.7524918.5–18.9
55Eikos-27Bv1.5 roster addendum A1Base model: Qwen/Qwen3.8-27Bsource18.519.518.5515018.3–18.6
56NInfer Qwen3.8-27B NVFP4 (T=1.5)detailsBase model: Qwen3.8-27Bsource18.518.918.5555118.2–18.6
57InstinctAPIdetailsBase model: Qwen3.8-27Bsource · Operator-reported; weights not publicly verifiable.18.318.918.3565218.0–18.5
58OpenJev (thinking, BF16)detailsBase model: undisclosedsource17.919.317.9545317.8–18.1
59djev (thinking)detailsBase model: google/diffusiongemma-26B-A4B-itsource17.418.517.4585617.3–17.5
60LitJevdetailsBase model: Qwen/Qwen3.8-27Bsource16.316.715.4605816.0–16.5
61Bev / Bonsai 27Bv1.5 roster addendum A4detailsBase model: Qwen/Qwen3.8-27Bsource15.816.012.3616215.1–16.0
62Raw Phi-4 mini direct logitsdetailsBase model: undisclosedsource15.213.110.6636310.3–20.7
63OpenSourceJevdetailsBase model: Qwen3.5-4Bsource13.211.29.264649.0–18.3
64reflex-27bdetailsBase model: Qwen3.8-27Bsource13.213.813.2626113.0–13.3
65Open-Jev 2BdetailsBase model: Qwen/Qwen3.5-2Bsource9.18.46.365666.6–11.5
66GLiNER2 largedetailsBase model: microsoft/deberta-v3-largesource8.47.25.867675.2–12.3
67Qwen3-Reranker-4BdetailsBase model: Qwen/Qwen3-4B-Basesource7.15.94.968684.3–10.7
68DeepSeek V4.1 FlashAPIdetailsBase model: undisclosed6.67.46.666656.6–6.7
69SimpleJevdetailsBase model: Qwen/Qwen3.5-0.8Bsource4.03.32.870702.1–6.7
70SimpleJev Qwen3.8-27BAPIdetailsBase model: Qwen/Qwen3.8-27Bsource3.43.73.469693.3–3.4
71decision-machine-1APIdetailsBase model: undisclosed3.22.52.371711.7–5.4
72GLiNER2.5 multidetailsBase model: mDeBERTa-v3-basesource2.72.11.972721.2–5.1
73Bosun v3.1 0.6Bv1.5 roster addendum A4detailsBase model: Qwen/Qwen3-0.6Bsource2.51.91.773731.2–4.6
74GLiNER2detailsBase model: DeBERTa-v3-basesource2.31.81.674740.9–4.5
75Deem 0.8B v1v1.5 roster addendum A4detailsBase model: Qwen/Qwen3.5-0.8Bsource2.11.71.575750.9–4.3
76JevActAPIdetailsBase model: undisclosedsource1.51.11.177760.5–3.2
77CLM-8BdetailsBase model: Qwen/Qwen3-8Bsource1.51.21.076770.5–3.2
78kev 0.6BdetailsBase model: Qwen/Qwen3-0.6B-Basesource1.30.90.978780.5–2.5
79Raw Qwen3 0.6B direct logitsdetailsBase model: Qwen/Qwen3-0.6B-Basesource1.10.90.779790.3–2.6
80GLiNER2.5 smalldetailsBase model: microsoft/deberta-v3-xsmallsource0.90.60.681800.2–2.2
81Raw Qwen3 1.7B direct logitsdetailsBase model: Qwen/Qwen3-1.7B-Basesource0.90.70.680810.1–2.3
82MirrordetailsBase model: undisclosed0.20.20.282820.0–0.9
83ZeroEntropy zerank-2detailsBase model: Qwen/Qwen3-4Bsource0.10.10.183830.0–0.5
84jeffdetailsBase model: undisclosedsource0.10.10.184840.0–0.5
85smalljev semantic-v9detailsBase model: openbmb/MiniCPM5-2B-Basesource0.10.00.085850.0–0.4
86Laya multilingualv1.5 roster addendum A4detailsBase model: mmBERT-basesource0.00.00.086860.0–0.3
87OpenDecisiondetailsBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.00.00.087870.0–0.2
88BAAI bge-reranker-v2-m3detailsBase model: BAAI/bge-m3source0.00.00.088880.0–0.0
89Certo v1detailsBase model: ModernBERT-largesource0.00.00.089890.0–0.0
90Decision FastdetailsBase model: Qwen/Qwen3-0.6B-Basesource0.00.00.090900.0–0.0
91Alibaba GTE Reranker ModernBERT-basedetailsBase model: answerdotai/ModernBERT-basesource0.00.00.091910.0–0.0
92kev 0.5BdetailsBase model: Qwen/Qwen2.5-0.5Bsource0.00.00.092920.0–0.0
93LayadetailsBase model: ModernBERT-largesource0.00.00.093930.0–0.0
94lev-350mdetailsBase model: LiquidAI/LFM2.5-350Msource0.00.00.094940.0–0.0
95Qwen3.5-0.8B Decision ModeldetailsBase model: Qwen/Qwen3.5-0.8B-Basesource0.00.00.095950.0–0.0
96Mixedbread mxbai-rerank-base-v2detailsBase model: undisclosedsource0.00.00.096960.0–0.0
97Needle 3detailsBase model: undisclosedsource0.00.00.097970.0–0.0
98Needle 3, options as toolsdetailsBase model: undisclosedsource0.00.00.098980.0–0.0
99open-jev-deberta-v3-largedetailsBase model: microsoft/deberta-v3-largesource0.00.00.099990.0–0.0
100Open Jev JSON CanvasdetailsBase model: google/diffusiongemma-26B-A4B-itsource0.00.00.01001000.0–0.0
101openJev VerdictdetailsBase model: undisclosedsource0.00.00.01011010.0–0.0
102openJev Verdict 1.4detailsBase model: knowledgator/gliclass-modern-base-v2.0source0.00.00.01021020.0–0.0
103Qwen3.8 27BAPIdetailsBase model: undisclosed0.00.00.01031030.0–0.0
104verdict-smalldetailsBase model: intfloat/multilingual-e5-smallsource0.00.00.01041040.0–0.0
105VondetailsBase model: answerdotai/ModernBERT-largesource0.00.00.01051050.0–0.0
106Laya typed-decisionsv1.5 roster addendum A4detailsBase model: ModernBERT-largesource0.00.00.01061060.0–0.0

What the run says

  • Among the 60 Jev-class systems, Jev 1.13.0 has the highest JevBench Capability Score, 80.0 (Intelligence 72.0, Calibration 88.0).
  • Cygnet leads the official JevBench v1.5.4 score (option A) with 73.7: Intelligence 71.1, Calibration 87.0, Speed 91.0, Cost 56.4 ($0.028 per 1,000 decisions).
  • The best open or open-planned rebuild, Winnow-12B Q8, is #2 at 73.2 — 0.5 points behind.
  • GPT-6 Luna (default medium reasoning effort) has the highest Intelligence (96.2) but places #35: Speed 73.2, Cost 38.3 — the harmonic mean does not let accuracy buy back a weak axis.
  • The strongest sealed Intelligence is 98.2 (Qwen3.8 27B, #103); sealed items carry half of Intelligence, and an open-minus-sealed gap beyond the field median plus eight points costs Intelligence.
  • 69 of 77 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant. No paired comparison is published for the other adjacent pairs, so no tie classification is inferred.
  • All 17 systems joined by separately hashed roster addenda and have official ranks in this revision; their A/B/C ranks and intervals are in the addendum table below.
  • classifier.dev scores 74.7 but is not ranked: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2).
  • SimpleJev Qwen3.6-35B-A3B, Decision-4B did not complete the full suite; they are listed without a rank.

Jev alternatives, open source and self-hosting

The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are Cygnet (#1, 73.7), Winnow-12B Q8 (#2, 73.2), JevK5 v0.3 (#4, 71.9), Plumb-4B (#5, 71.6). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost — 25/25/25/25 — with an Intelligence floor of 50 and low-axis gates on Speed and Cost. Version v1.5.4 measures 904 open and 720 sealed decisions per system; Choice, Noul and Score each carry a third, sealed items contribute 50% of Intelligence, and an open-minus-sealed gap beyond the field median costs points. Method notes · options B and C.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version or a disclosed roster addendum. For private data, see the custom evaluation options.

What a decision costs

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole typed request — state, rubric and options — not a single token.

How costs are estimated

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured — 3 rows carry a tariff. Systems without one — open weights, author demos, models we ran ourselves — are priced as if a large inference provider hosted them: the list price of the same weights, or the nearest larger sibling or size class when the exact weights are not listed. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

Price rules (v1.5). Only public, bookable list prices that have been in effect for at least 30 days count; a manufacturer's standard, non-promotional launch list price counts from day one, and promotions, subsidies, credits and free tiers never do. The scoring price is never below the market reference price of the system's base model. A system without any eligible price is listed as unpriced — no Cost axis and no score until a price qualifies. A later price change triggers a re-score with a visible note on the row.

API models with a known base model (from 1 Oct 2026). They are ranked at their developer's own stated API price; a striped second bar shows the score and rank they would have at the base-model reference price, the way we price self-served open weights of the same base. First applied on Image JevBench v0.1.5 (Wity-1). The v1.5.4 scores on this page are unchanged; JevBench rows follow this rule from the next version.

  • Cygnet — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Winnow-12B Q8 — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Jev 1.13.0 — ~$0.032 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • Jev-Omni — ~$0.029 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • decider-4b v2 — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • SemIf — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • spark-s1-4b-v6 — ~$0.021 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • metask-jev-4b — ~$0.026 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Hopper — ~$0.018 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Malkuth-4B — ~$0.030 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • reflex 4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • jev-local — ~$0.024 est. per 1,000 decisions: price floor: base-model reference price applied
  • djev (Maisa, diffusion-gemma) — ~$0.053 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Raw Qwen3 4B Instruct 2507 direct logits — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • jqv — ~$0.042 est. per 1,000 decisions: documented hosted-model estimate
  • JevK5 v0.2.0 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Qwen3.5-9B Jev-like data-mix v2 — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
  • Standard One 8B — ~$0.078 est. per 1,000 decisions: documented hosted-model estimate
  • NInfer Qwen3.8-Flash-Next mixed — ~$0.082 est. per 1,000 decisions: documented hosted-model estimate
  • swanOne — ~$0.085 est. per 1,000 decisions: price floor: base-model reference price applied
  • Raw Qwen3 8B direct logits — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
  • decider-2b — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate
  • system-one — ~$0.068 est. per 1,000 decisions: documented hosted-model estimate
  • system-one-open — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • GPT-6 Luna (low reasoning effort) — ~$0.108 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • GPT-6 Luna (default medium reasoning effort) — ~$0.114 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • JevOne — ~$0.101 est. per 1,000 decisions: documented hosted-model estimate
  • kev 4B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • kev 8B — ~$0.097 est. per 1,000 decisions: price floor: base-model reference price applied
  • open-alternative-jev — ~$0.017 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Bespoke Nimble 9B — ~$0.128 est. per 1,000 decisions: documented hosted-model estimate
  • Malkuth-2B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • openjev-sglang — ~$0.140 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • decider-35b-a3b — ~$0.154 est. per 1,000 decisions: price floor: base-model reference price applied
  • local-jev Qwen3.5-4B — ~$0.023 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Open-Jev 9B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
  • Decision 2B — ~$0.013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • typecastlm — ~$0.016 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • JEV Qwen3.5-9B Base NVFP4 — ~$0.056 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference
  • NInfer Qwen3.8-27B NVFP4 — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
  • NInfer Qwen3.8-27B NVFP4 (T=1.5) — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
  • Instinct — ~$0.230 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • OpenJev (thinking, BF16) — ~$0.241 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • djev (thinking) — ~$0.249 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • LitJev — ~$0.244 est. per 1,000 decisions: price floor: base-model reference price applied
  • Raw Phi-4 mini direct logits — ~$0.036 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • OpenSourceJev — ~$0.011 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • reflex-27b — ~$0.297 est. per 1,000 decisions: price floor: base-model reference price applied
  • Open-Jev 2B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
  • GLiNER2 large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Qwen3-Reranker-4B — ~$0.052 est. per 1,000 decisions: ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
  • DeepSeek V4.1 Flash — ~$0.498 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • SimpleJev — ~$0.0098 est. per 1,000 decisions: documented hosted-model estimate
  • SimpleJev Qwen3.8-27B — ~$0.687 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • decision-machine-1 — ~$0.029 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • GLiNER2.5 multi — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • GLiNER2 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • JevAct — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • CLM-8B — ~$0.045 est. per 1,000 decisions: price floor: base-model reference price applied
  • kev 0.6B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Raw Qwen3 0.6B direct logits — ~$0.0056 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • GLiNER2.5 small — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Raw Qwen3 1.7B direct logits — ~$0.011 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Mirror — ~$0.0023 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • ZeroEntropy zerank-2 — ~$0.052 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
  • jeff — ~$0.0043 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • smalljev semantic-v9 — ~$0.020 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • OpenDecision — ~$0.0050 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • BAAI bge-reranker-v2-m3 — ~$0.023 est. per 1,000 decisions: ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3)
  • Certo v1 — ~$0.0013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Decision Fast — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Alibaba GTE Reranker ModernBERT-base — ~$0.011 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies
  • kev 0.5B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Laya — ~$0.0032 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • lev-350m — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Qwen3.5-0.8B Decision Model — ~$0.0048 est. per 1,000 decisions: documented hosted-model estimate
  • Mixedbread mxbai-rerank-base-v2 — ~$0.021 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies
  • Needle 3 — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Needle 3, options as tools — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • open-jev-deberta-v3-large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Open Jev JSON Canvas — ~$0.049 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • openJev Verdict — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • openJev Verdict 1.4 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Qwen3.8 27B — ~$2.18 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • verdict-small — ~$0.0009 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Von — ~$0.0038 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • classifier.dev — ~$0.023 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies
  • JevK5 v0.3 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Plumb-4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Decision 4B v1.2 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Imajev-4B — ~$0.017 est. per 1,000 decisions: ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B.
  • Decision 4B v1.1 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Surogate Rune 26B-A4B v3 — ~$0.050 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
  • AutoJev-27B (denis-pplx, Qwen3.8-27B) — ~$0.226 est. per 1,000 decisions: documented hosted-model estimate
  • AutoJev-27B (RTX PRO 6000) — ~$0.226 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
  • Eikos-27B — ~$0.238 est. per 1,000 decisions: documented hosted-model estimate
  • SimpleJev Qwen3.6-35B-A3B — ~$0.145 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • Nemotron Diffusion 8B — ~$0.030 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Bev / Bonsai 27B — ~$0.247 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b.
  • Deem 0.8B v1 — ~$0.0045 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B.
  • Decision-4B — ~$0.014 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.
  • Instinct Dual 4B — ~$0.021 est. per 1,000 decisions: ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.
  • Bosun v3.1 0.6B — ~$0.0056 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate.
  • Laya multilingual — ~$0.0039 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.
  • Laya typed-decisions — ~$0.0038 est. per 1,000 decisions: ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.
  • Manchego v2.1 — ~$0.015 est. per 1,000 decisions: ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes.

Roster addendum: newcomers scored on the same frozen protocol (17)

All 17 A1/A2/A3/A4/A5 systems below completed all 1,624 decisions and now have official ranks in A, B, and C; the interactive score presets also include them. Their scores, individual 95% intervals, and the frozen v1.5.0 G_med are unchanged. No new paired-bootstrap comparisons were computed for addendum systems. Existing tie markers are retained only for base-system pairs that remain adjacent; no tie or separation is inferred for the other pairs.

SystemRank (A)A score · 95% CIRank (B)B score · 95% CIRank (C)C score · 95% CI
JevK5 v0.3v1.5 roster addendum A1#471.9 69.4–72.9#568.1 64.9–69.8#863.2 51.3–72.1
Plumb-4Bv1.5 roster addendum A1#571.6 69.2–72.7#667.7 64.7–69.6#962.0 50.7–70.9
Decision 4B v1.2v1.5 roster addendum A1#870.8 68.5–72.0#966.6 63.7–68.4#1256.7 47.6–65.1
Imajev-4Bv1.5 roster addendum A2#970.4 67.8–71.6#1166.2 63.1–68.1#1355.9 47.1–64.5
Decision 4B v1.1v1.5 roster addendum A1#1070.4 66.9–71.6#1266.1 62.1–67.9#1455.1 46.4–63.5
Manchego v2.1v1.5 roster addendum A5#1168.8 59.6–70.3#1464.4 55.2–66.5#2050.1 41.4–58.1
Surogate Rune 26B-A4B v3v1.5 roster addendum A2#1766.5 65.4–67.0#1066.6 65.1–67.5#666.5 65.4–67.0
Instinct Dual 4Bv1.5 roster addendum A4API#2747.0 37.9–56.8#3143.0 34.2–52.8#3232.6 26.3–39.4
Nemotron Diffusion 8Bv1.5 roster addendum A3#4525.7 18.6–32.8#4722.7 16.1–29.5#5517.8 12.9–22.8
AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1#5219.5 19.3–19.6#4920.4 20.1–20.6#4719.5 19.3–19.6
AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2#5319.5 19.3–19.6#5020.4 20.1–20.6#4819.5 19.3–19.6
Eikos-27Bv1.5 roster addendum A1#5518.5 18.3–18.6#5119.5 19.2–19.7#5018.5 18.3–18.6
Bev / Bonsai 27Bv1.5 roster addendum A4#6115.8 15.1–16.0#6116.0 15.2–16.4#6212.3 10.5–14.3
Bosun v3.1 0.6Bv1.5 roster addendum A4#732.5 1.2–4.6#731.9 0.9–3.6#731.7 0.8–3.2
Deem 0.8B v1v1.5 roster addendum A4#752.1 0.9–4.3#751.7 0.7–3.5#751.5 0.6–3.0
Laya multilingualv1.5 roster addendum A4#860.0 0.0–0.3#860.0 0.0–0.2#860.0 0.0–0.2
Laya typed-decisionsv1.5 roster addendum A4#1060.0 0.0–0.0#1060.0 0.0–0.0#1060.0 0.0–0.0

Ranks follow option scores; score intervals are per system, not pairwise rank comparisons. Every slider preset sorts these same ranked systems using its selected weights.

Not ranked: partial, unpriced and unmeasured systems

These systems are part of the 112-system v1.5 roster but have no rank. Their numbers are never shown as zero or free.

Partial runs (2)

  • SimpleJev Qwen3.6-35B-A3BAPIpartial run: Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.
  • Decision-4B (Eval Engine / Chromia)v1.5 roster addendum A4not ranked: PARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.

Incomplete or not measured in v1.5 (3)

  • Jobe Qwen3.5-4B (frozen): not measured. Base model: undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.
  • mica-v01-4bv1.5 roster addendum A1: The frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank. (1,088/1,624 rows; 536 missing). Base model: undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.
  • OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16): not measured. Base model: undisclosed · No verified public base-model disclosure recorded for this unmeasured evaluated variant.

Incomplete and unmeasured systems receive no official rank. Existing results from earlier benchmark versions remain on their frozen version pages.

Method notes: what changed in v1.5

The four axes

Intelligence · 25%
How often answers are right above chance: each type is normalized against its task-specific random baseline (chance = 0, perfect = 100; below-chance tiers can be negative). Choice, Noul and Score count one third each. Easy, Standard, Judge and Hard items count 10%, 20%, 30% and 40%; the open and sealed sets count equally.
Calibration · 25%
How closely stated probabilities match what happens. It uses ECE and TVD for Choice, ECE and Brier for Noul, and normalized RPS plus top-level ECE for Score. The three types count equally; open and sealed items are pooled.
Speed · 25%
Serial response latency on open Standard and Judge items. The p50 and p95 each get a log score: 100 − 20 × log₁₀(seconds ÷ 0.1), then are averaged. Self-hosted and demo endpoints get the published ×2 plus 0.15-second adjustment.
Cost · 25%
Estimated or billed US dollars per 1,000 decisions, using pooled token use across 1,624 decisions and the documented price rules. The log score is 100 − 30 × log₁₀(cost ÷ $0.001). The price reference is $0.001 per 1,000 decisions.

The official score is the weighted harmonic mean of the four axes. Option A gives each axis 25%; option B (40/20/20/20) is a secondary view, while option C keeps equal axes and sets the Intelligence gate at 60. In A and B, Intelligence, Speed and Cost each have a quadratic gate below 50. To limit benchmaxxing on the public items, the open-minus-sealed Intelligence gap may be up to 8 points above the field median (G_med) before a penalty applies. Each further point lowers the multiplier on unpenalized Intelligence by one percentage point.

Frozen method METHOD-v1.5, SHA-256 c25d3d8b8512…; pricing addendum v1.5-M2, SHA-256 2fc44459ef80….

The method owner chose equal axis weights and equal weights for Choice, Noul and Score after reviewing the What-If Lab, preserving continuity with v1.4 and treating the three decision types equally. Disclosed headline amendment: equal-axis, equal-type A, SHA-256 752ddccc4e19…. B remains a secondary view.

Re-evaluation policy. Every release re-evaluates the current top 10 on the composite score. Models ranked #11 and below are re-evaluated on a slower cadence — at least monthly, or with every third scheduled refresh release, whichever comes first — and their score is shown as last measured on its release. A material method change re-evaluates every model. Paid fast-lane runs are evaluated within 48 hours of payment, and new submissions are evaluated in the order received.

  • 1,624 decisions per system: 904 open (601 published) and 720 sealed, drawn fresh from a private pool with the same tier mix as the open set. Sealed counts for 50% of Intelligence: base = 0.5 × I_open + 0.5 × I_sealed.
  • Three request types are scored natively and chance-corrected per item: Choice, Noul and Score each receive one third. Tier weights easy / standard / judge / hard = 10 / 20 / 30 / 40. A type a system does not support is excluded, never scored zero; only full-coverage systems are ranked.
  • Overfit penalty relative to the field: excess = gap − G_med, penalty = max(0, 1 − max(0, excess − 8) / 100). G_med for this batch is 5.2 CC points.
  • Calibration is typed (Choice ECE/TVD, Noul ECE with Brier, Score normalised RPS and top-level ECE), pooled over open and sealed. Speed and Cost formulas are unchanged from v1.4; self-hosted and demo endpoints carry the ×2 + 0.15 s adjustment. A manufacturer's standard, non-promotional launch list price counts from day one, but a newer price cut younger than 30 days does not. Rows without token counts use the measured proxy-token basis. A system without any eligible public, bookable price is listed as unpriced.
  • The frozen 25 Sep DeepInfra snapshot records Qwen3.5-4B as deprecated on 11 Jun 2026 and replaced by Qwen3.5-9B. Its frozen snapshot rates remain the v1.5 M2 reference; price basis tooltips and the correction note disclose this. Pricing disclosure correction SHA-256: 1b660648bd49….
  • Composite: weighted harmonic mean with the Intelligence, Speed and Cost gates below 50 (Intelligence below 60 in option C). The official headline A uses equal 25 / 25 / 25 / 25 axis weights and Intelligence floor 50. B remains the secondary 40 / 20 / 20 / 20 view; C keeps equal axes and Intelligence floor 60. Ties come from the paired bootstrap. The gates belong to the score, not to the weights: in the custom-weight views of the score chart they still apply to an axis set to 0, as in the published views — so “Intelligence only” follows the Intelligence column except for systems with Cost, Speed or Intelligence below 50 (a general-purpose LLM with Cost 39 keeps only (39/50)², about 0.61, of its score). Every gated row carries a “gate ×…” tag naming the axis and factor, and the weights panel lists the gates that fire.
  • All 17 full-coverage systems marked with a v1.5 roster addendum label are officially ranked in A, B, C, and every score preset. Their scores, individual intervals, method, pricing rules and frozen G_med are unchanged. No new paired-bootstrap comparison is inferred for addendum pairs.
  • Before every release we review the leaderboard for anomalies and close loopholes with general, documented rules. The page and Git repository provide transparent data and method details; Benchmark Heaven owns its rules.

Data file SHA-256 0cf210b76bf85084a5f3fb40fb109e9a2c2f93df42ff6628696377666e89db45 · scorer output SHA-256 452885de2a84cd5b9ed393d541fd8d6c6a540f9d7f762f9d2df9349383f26173 · run kind official.

Limits
  • 1,624 decisions per system (904 open, 720 sealed) is a measurement, not a census, and it is English-only.
  • The weights are a choice. Option A weights the four axes equally and uses a harmonic mean, so the weakest axis dominates; options B and C are published alternatives and the weight sliders re-score the same axes for exploration — only the official option gives the official score and rank. If a wrong decision costs you more than a slow or expensive one, read the Intelligence column and the per-type competence rather than the score alone.
  • The latency adjustment (×2, +0.15 s on our own servers and demo endpoints) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. Serving under load trades per-user speed for throughput. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo.
  • Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
  • Latency is one origin at one time of day; hosted endpoints, public demos and our own pods are different kinds of latency. Public demo endpoints are shared with everyone else using them.
  • Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version or a disclosed roster addendum rather than silently changing this one.

3D view: three.js r128 (MIT).

Loading capability views…

Loading context-length views…

Previous release: JevBench v1.5.3 (frozen results).

Model author? Request a priority evaluation · Submit a model for evaluation →

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.