JevBench v1.5.0 — Jev alternatives ranking

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Frozen release · 1,624 decisions per system (904 open + 720 sealed; sealed decisions are half of Intelligence) · 89 ranked of 103 roster systems · only system-level sealed aggregates are published · aggregate results JSON sha256 6b2f6b058b36…

Making decisions from images? Explore Image JevBench v0.1.4 and compare its systems.

Share this version · View live board · Previous release: JevBench v1.4.2.2

JevBench v1.5.0 · headline ranking

Capability ranking of Jev-class systems

Capability averages Intelligence and Calibration. classifier.dev leads the Jev-class systems with 82.5.

Jev-class means at most 2× Jev's cost and median latency. How we choose ↘

#SystemICCap.$/1k
  1. –classifier.dev75.858.982.5$0.023*Not ranked in the official JevBench Score (honorable mention)
  2. 1Jev 1.13.072.054.780.0$0.032*Reference system for the Jev-class limits
  3. 2Winnow-12B Q874.456.679.3$0.028*
  4. 3Cygnet71.156.479.0$0.028*
  5. –Surogate Rune 26B-A4B v369.749.079.0$0.050*Not ranked in the official JevBench Score (addendum)
  6. 4Jev-Omni70.556.176.5$0.029*
  7. 5djev72.348.276.4$0.053*
  8. –JevK5 v0.356.363.172.3$0.017*Not ranked in the official JevBench Score (addendum)
  9. –Plumb-4B55.863.171.6$0.017*Not ranked in the official JevBench Score (addendum)
  10. –Decision 4B v1.253.763.171.1$0.017*Not ranked in the official JevBench Score (addendum)
Show all 56 Jev-class systems (46 more)
  1. –Imajev-4B53.563.370.8$0.017*Not ranked in the official JevBench Score (addendum)
  2. 6decider-4b v255.864.570.7$0.015*
  3. –Decision 4B v1.153.163.170.2$0.017*Not ranked in the official JevBench Score (addendum)
  4. 7Hopper49.962.368.9$0.018*
  5. 8Malkuth-4B54.555.868.8$0.030*
  6. 9metask-jev-4b53.557.768.1$0.026*
  7. 10jqv49.151.267.9$0.042*
  8. 11SemIf51.363.167.6$0.017*
  9. 12jev-local56.358.967.0$0.024*
  10. 13spark-s1-4b-v662.160.465.9$0.021*
  11. 14JevK5 v0.2.046.763.165.8$0.017*
  12. 15Decision 2B31.366.458.8$0.013*
  13. 16local-jev Qwen3.5-4B33.759.358.1$0.023*
  14. 17open-alternative-jev37.463.357.1$0.017*
  15. 18decider-2b42.364.956.9$0.015*
  16. 19system-one-open41.668.356.6$0.011*
  17. 20JEV Qwen3.5-9B Base NVFP432.247.656.6$0.056*
  18. 21Malkuth-2B35.565.655.4$0.014*
  19. 22typecastlm31.264.054.0$0.016*
  20. 23kev 4B39.965.853.8$0.014*
  21. 24Raw Qwen3 4B Instruct 2507 direct logits54.163.453.5$0.017*
  22. 25Raw Phi-4 mini direct logits27.653.249.5$0.036*
  23. 26Qwen3-Reranker-4B21.048.448.6$0.052*
  24. 27decision-machine-114.856.348.0$0.029*
  25. 28Certo v10.096.943.8$0.0013*
  26. 29Mixedbread mxbai-rerank-base-v20.060.343.3$0.021*
  27. 30ZeroEntropy zerank-24.848.443.2$0.052*
  28. 31Von0.082.741.7$0.0038*
  29. 32BAAI bge-reranker-v2-m30.059.341.7$0.023*
  30. 33openJev Verdict 1.40.086.640.2$0.0028*
  31. 34kev 0.6B10.380.139.0$0.0046*
  32. 35lev-350m0.080.038.8$0.0046*
  33. 36Open Jev JSON Canvas77.149.338.6$0.049*
  34. 37smalljev semantic-v93.760.838.4$0.020*
  35. 38Decision Fast0.080.138.0$0.0046*
  36. 39Alibaba GTE Reranker ModernBERT-base0.069.337.6$0.011*
  37. 40OpenDecision2.179.137.3$0.0050*
  38. 41JevAct11.268.636.9$0.011*
  39. 42GLiNER2.5 small9.286.632.5$0.0028*
  40. 43kev 0.5B0.080.132.5$0.0046*
  41. 44CLM-8B11.450.530.0$0.045*
  42. 45verdict-small0.0100.029.7$0.00087*
  43. 46openJev Verdict0.086.626.1$0.0028*
  44. 47GLiNER213.586.624.5$0.0028*
  45. 48Raw Qwen3 0.6B direct logits10.677.616.0$0.0056*
  46. 49Raw Qwen3 1.7B direct logits9.768.515.6$0.011*

Wide coloured bar = Capability (0–100). Thin red line = cost per 1,000 decisions; log scale, each gridline = 10×, shorter is cheaper. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ for the full values, median latency and cost relative to Jev.

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
  • Cost line
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability, not numbered. Each row says which limit it misses, measured against Jev 1.13.0 ($0.032 per 1,000 decisions, median 0.62 s).

  1. –Qwen3.8 27B95.60.096.8$2.18*Outside: cost 67.4× Jev, latency 10.5× Jev
  2. –GPT-6 Luna (medium)96.238.395.9$0.11*Outside: cost 3.5× Jev, latency 2.5× Jev
  3. –DeepSeek V4.1 Flash93.719.195.3$0.50*Outside: cost 15.4× Jev, latency 2.9× Jev
  4. –GPT-6 Luna (low)95.339.195.1$0.11*Outside: cost 3.3× Jev, latency 2.6× Jev
  5. –GPT-5.6 Luna94.330.794.5$0.20Outside: cost 6.3× Jev, latency 2.1× Jev
  6. –djev77.328.186.5$0.25*Outside: cost 7.7× Jev, latency 2.4× Jev
  7. –OpenJev84.228.583.7$0.24*Outside: cost 7.5× Jev, latency 2.5× Jev
  8. –Autoloops – Gemma 4 31B IT76.739.681.3$0.10Outside: cost 3.2× Jev
  9. –Eikos-27B75.128.780.7$0.24*Outside: cost 7.4× Jev
  10. –AutoJev-27B72.829.480.2$0.23*Outside: cost 7.0× Jev
  11. –SimpleJev Qwen3.8-27B72.814.980.0$0.69*Outside: cost 21.3× Jev, latency 2.7× Jev
  12. –AutoJev-27B72.829.479.7$0.23*Outside: cost 7.0× Jev
  13. –swanOne71.242.279.1$0.085*Outside: cost 2.6× Jev
  14. –NInfer Qwen3.8-Flash-Next mixed67.242.577.9$0.082*Outside: cost 2.5× Jev
  15. –Gemini 3.1 Flash-Lite77.629.876.1$0.22Outside: cost 6.8× Jev
  16. –NInfer Qwen3.8-27B NVFP465.529.175.7$0.23*Outside: cost 7.1× Jev
  17. –reflex-27b62.825.874.4$0.30*Outside: cost 9.2× Jev, latency 4.5× Jev
  18. –Instinct62.729.273.9$0.23*Outside: cost 7.1× Jev
  19. –NInfer Qwen3.8-27B NVFP461.129.173.8$0.23*Outside: cost 7.1× Jev
  20. –Open-Jev 9B63.833.172.7$0.17*Outside: cost 5.3× Jev, latency 2.1× Jev
  21. –LitJev58.328.471.4$0.24*Outside: cost 7.6× Jev, latency 5.0× Jev
  22. –Standard One 8B59.643.371.4$0.078*Outside: cost 2.4× Jev
  23. –decider-35b-a3b60.534.471.2$0.15*Outside: cost 4.8× Jev
  24. –openjev-sglang58.635.670.8$0.14*Outside: cost 4.3× Jev
  25. –Qwen3.5-9B Jev-like data-mix v260.445.770.7$0.065*Outside: cost 2.0× Jev
  26. –Bespoke Nimble 9B63.736.870.5$0.13*Outside: cost 4.0× Jev
  27. –JevOne54.139.869.5$0.10*Outside: cost 3.1× Jev
  28. –reflex 4B51.563.169.2$0.017*Outside: latency 4.6× Jev
  29. –kev 8B48.340.459.7$0.097*Outside: cost 3.0× Jev
  30. –Open-Jev 2B33.633.153.6$0.17*Outside: cost 5.3× Jev
  31. –OpenSourceJev25.869.150.9$0.011*Outside: latency 2.3× Jev
  32. –Raw Qwen3 8B direct logits51.145.550.2$0.065*Outside: cost 2.0× Jev
  33. –system-one50.545.049.9$0.068*Outside: cost 2.1× Jev
  34. –jeff4.481.142.4$0.0043*Outside: latency 11.6× Jev
  35. –SimpleJev Qwen3.6-35B-A3B0.035.139.2$0.15*Outside: cost 4.5× Jev, latency 2.7× Jev
  36. –open-jev-deberta-v3-large0.077.638.5$0.0056*Outside: latency 4.8× Jev
  37. –Laya0.084.936.9$0.0032*Outside: latency 2.4× Jev
  38. –Qwen3.5-0.8B Decision Model0.079.536.8$0.0048*Outside: latency 2.1× Jev
  39. –GLiNER2.5 multi14.086.636.0$0.0028*Outside: latency 2.2× Jev
  40. –GLiNER2 large22.577.632.4$0.0056*Outside: latency 3.0× Jev
  41. –SimpleJev16.770.331.7$0.0098*Outside: latency 12.0× Jev
  42. –Mirror5.689.324.4$0.0023*Outside: latency 7.4× Jev
  43. –Needle 30.061.60.0$0.019*Outside: latency 219.8× Jev
  44. –Needle 3, options as tools0.061.60.0$0.019*Outside: latency 94.3× Jev

Jev-class = cost per decision at most 2× Jev 1.13.0's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev 1.13.0's (≤ 1.23 s, the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score). 56 of 100 systems qualify; the other 44, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability; the official JevBench Score weighs all four axes.

Capability against cost and speed

Jev-class systems are shown by default. Bubble size follows the official JevBench Score. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 56 systems. Upper right is best: more capable and cheaper.102030405060708090100$0.0010$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev← priciercheaper →1. Jev 1.13.02. Winnow-12B Q83. Cygnet4. Jev-Omni5. djev
56 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 56 systems. Upper right is best: more capable and faster.10203040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev latency← slowerfaster →1. Jev 1.13.02. Winnow-12B Q83. Cygnet4. Jev-Omni5. djev
56 systems. Tap a bubble for its values.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
  • faint = outside Jev-class

What-If: the weight sliders below re-score every system under other axis weights — only the equal 25/25/25/25 weights give the official option-A ranking. The 3D view of capability, cost and speed loads further down.

JevBench v1.5.0

JevBench Composite Score: 89 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

Weights:
Adjust weights ↓
View by:Capability ↑

Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).

Whiskers are 95% bootstrap intervals. 75 of 88 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant.

Greener = stronger within its column.

100 of 100 systems, sorted by official rank, #1 first.

  1. 1Cygnet73.7I 71C 87S 91K 56est.$0.028
  2. 2Winnow-12B Q873.2I 74C 84S 86K 57est.$0.028
  3. 3Jev 1.13.0API72.1I 72C 88S 84K 55est.$0.032
  4. 4Jev-Omni71.5I 70C 83S 85K 56est.$0.029
  5. 5decider-4b v271.3I 56C 86S 91K 65est.$0.015
  6. 6SemIf68.7I 51C 84S 91K 63est.$0.017
  7. 7spark-s1-4b-v668.2I 62C 70S 86K 60est.$0.021
  8. 8metask-jev-4b67.5I 53C 83S 89K 58est.$0.026
  9. 9Hopper67.5I 50C 88S 87K 62est.$0.018
  10. 10Malkuth-4B66.8I 54C 83S 86K 56est.$0.030
  11. 11reflex 4B65.2I 51C 87S 69K 63est.$0.017
  12. 12jev-local65.2I 56C 78S 73K 59est.$0.024
  13. 13djev (Maisa, diffusion-gemma)64.2I 72C 80S 91K 48est.$0.053
  14. 14Raw Qwen3 4B Instruct 2507 direct logits62.1I 54C 53S 89K 63est.$0.017
  15. 15jqv60.8I 49C 87S 83K 51est.$0.042
  16. 16JevK5 v0.2.058.1I 47C 85S 91K 63est.$0.017
  17. 17Qwen3.5-9B Jev-like data-mix v253.0I 60C 81S 82K 46est.$0.065
  18. 18Standard One 8B47.8I 60C 83S 92K 43est.$0.078
  19. 19NInfer Qwen3.8-Flash-Next mixed47.5I 67C 89S 89K 43est.$0.082
  20. 20swanOne46.6I 71C 87S 85K 42est.$0.085
Show all 100 systems (69 more ranked, 11 more not ranked)
  1. 21Raw Qwen3 8B direct logits45.2I 51C 49S 87K 46est.$0.065
  2. 22decider-2b45.1I 42C 72S 94K 65est.$0.015
  3. 23system-one44.1I 50C 49S 91K 45est.$0.068
  4. 24system-one-openAPI42.4I 42C 72S 78K 68est.$0.011
  5. 25Autoloops – Gemma 4 31B ITAPI40.5I 77C 86S 84K 40$0.103
  6. 26GPT-6 Luna (low reasoning effort)API40.5I 95C 95S 73K 39est.$0.108
  7. 27GPT-6 Luna (default medium reasoning effort)API38.8I 96C 96S 73K 38est.$0.114
  8. 28JevOne38.2I 54C 85S 90K 40est.$0.101
  9. 29kev 4B38.1I 40C 68S 85K 66est.$0.014
  10. 30kev 8B34.2I 48C 71S 84K 40est.$0.097
  11. 31open-alternative-jev33.6I 37C 77S 91K 63est.$0.017
  12. 32Bespoke Nimble 9B31.8I 64C 77S 83K 37est.$0.128
  13. 33Malkuth-2B29.9I 36C 75S 92K 66est.$0.014
  14. 34openjev-sglangAPI29.0I 59C 83S 78K 36est.$0.140
  15. 35decider-35b-a3b27.5I 60C 82S 91K 34est.$0.154
  16. 36local-jev Qwen3.5-4B25.8I 34C 82S 84K 59est.$0.023
  17. 37Open-Jev 9B24.4I 64C 82S 74K 33est.$0.170
  18. 38Decision 2B22.5I 31C 86S 90K 66est.$0.013
  19. 39GPT-5.6 LunaAPI22.4I 94C 95S 74K 31$0.205
  20. 40typecastlm21.8I 31C 77S 92K 64est.$0.016
  21. 41JEV Qwen3.5-9B Base NVFP420.1I 32C 81S 94K 48est.$0.056
  22. 42Gemini 3.1 Flash-LiteAPI19.6I 78C 75S 80K 30$0.219
  23. 43NInfer Qwen3.8-27B NVFP418.7I 65C 86S 90K 29est.$0.231
  24. 44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.5I 61C 86S 90K 29est.$0.231
  25. 45InstinctAPI18.3I 63C 85S 82K 29est.$0.230
  26. 46OpenJev17.9I 84C 83S 74K 29est.$0.241
  27. 47djev (thinking)17.4I 77C 96S 72K 28est.$0.249
  28. 48LitJev16.3I 58C 84S 68K 28est.$0.244
  29. 49Raw Phi-4 mini direct logits15.2I 28C 71S 89K 53est.$0.036
  30. 50OpenSourceJev13.2I 26C 76S 73K 69est.$0.011
  31. 51reflex-27b13.2I 63C 86S 69K 26est.$0.297
  32. 52Open-Jev 2B9.1I 34C 74S 76K 33est.$0.170
  33. 53GLiNER2 large8.4I 22C 42S 65K 78est.$0.0056
  34. 54Qwen3-Reranker-4B7.1I 21C 76S 80K 48est.$0.052
  35. 55DeepSeek V4.1 FlashAPI6.6I 94C 97S 69K 19est.$0.498
  36. 56SimpleJev4.0I 17C 47S 59K 70est.$0.0098
  37. 57SimpleJev Qwen3.8-27BAPI3.4I 73C 87S 75K 15est.$0.687
  38. 58decision-machine-1API3.2I 15C 81S 93K 56est.$0.029
  39. 59GLiNER2.5 multi2.7I 14C 58S 67K 87est.$0.0028
  40. 60GLiNER22.3I 13C 36S 70K 87est.$0.0028
  41. 61JevActAPI1.5I 11C 63S 76K 69est.$0.011
  42. 62CLM-8B1.5I 11C 48S 93K 51est.$0.045
  43. 63kev 0.6B1.3I 10C 68S 87K 80est.$0.0046
  44. 64Raw Qwen3 0.6B direct logits1.1I 11C 21S 90K 78est.$0.0056
  45. 65GLiNER2.5 small0.9I 9C 56S 77K 87est.$0.0028
  46. 66Raw Qwen3 1.7B direct logits0.9I 10C 22S 90K 69est.$0.011
  47. 67Mirror0.2I 6C 43S 64K 89est.$0.0023
  48. 68ZeroEntropy zerank-20.1I 5C 82S 80K 48est.$0.052
  49. 69jeff0.1I 4C 80S 56K 81est.$0.0043
  50. 70smalljev semantic-v90.1I 4C 73S 86K 61est.$0.020
  51. 71OpenDecision0.0I 2C 73S 87K 79est.$0.0050
  52. 72BAAI bge-reranker-v2-m30.0I 0C 83S 91K 59est.$0.023
  53. 73Certo v10.0I 0C 88S 91K 97est.$0.0013
  54. 74Decision Fast0.0I 0C 76S 91K 80est.$0.0046
  55. 75Alibaba GTE Reranker ModernBERT-base0.0I 0C 75S 91K 69est.$0.011
  56. 76kev 0.5B0.0I 0C 65S 88K 80est.$0.0046
  57. 77Laya0.0I 0C 74S 74K 85est.$0.0032
  58. 78lev-350m0.0I 0C 78S 94K 80est.$0.0046
  59. 79Qwen3.5-0.8B Decision Model0.0I 0C 74S 72K 80est.$0.0048
  60. 80Mixedbread mxbai-rerank-base-v20.0I 0C 87S 89K 60est.$0.021
  61. 81Needle 30.0I 0C 0S 34K 62est.$0.019
  62. 82Needle 3, options as tools0.0I 0C 0S 41K 62est.$0.019
  63. 83open-jev-deberta-v3-large0.0I 0C 77S 68K 78est.$0.0056
  64. 84Open Jev JSON Canvas0.0I 77C 0S 86K 49est.$0.049
  65. 85openJev Verdict0.0I 0C 52S 84K 87est.$0.0028
  66. 86openJev Verdict 1.40.0I 0C 80S 81K 87est.$0.0028
  67. 87Qwen3.8 27BAPI0.0I 96C 98S 57K 0est.$2.178
  68. 88verdict-small0.0I 0C 59S 82K 100est.$0.0009
  69. 89Von0.0I 0C 83S 76K 83est.$0.0038
  70. classifier.dev (honorable mention)API74.7I 76C 89S 82K 59est.$0.023
  71. JevK5 v0.3 (addendum)new71.9I 56C 88S 94K 63est.$0.017
  72. Plumb-4B (addendum)71.6I 56C 87S 93K 63est.$0.017
  73. Decision 4B v1.2 (addendum)new70.8I 54C 89S 94K 63est.$0.017
  74. Imajev-4B (addendum)new70.4I 53C 88S 91K 63est.$0.017
  75. Decision 4B v1.1 (addendum)new70.4I 53C 87S 94K 63est.$0.017
  76. Surogate Rune 26B-A4B v3 (addendum)new66.5I 70C 88S 86K 49est.$0.050
  77. AutoJev-27B (denis-pplx, Qwen3.8-27B) (addendum)new19.5I 73C 88S 88K 29est.$0.226
  78. AutoJev-27B (RTX PRO 6000) (addendum)new19.5I 73C 87S 87K 29est.$0.226
  79. Eikos-27B (addendum)new18.5I 75C 86S 88K 29est.$0.238
  80. SimpleJev Qwen3.6-35B-A3B (partial run)API0.0I 0C 78S 75K 35est.$0.145
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • system-one-open
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers; new = first listed in v1.5.0; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

Compare two systems

Pick any two. Four radars: the score axes, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 72.1 (#3)
  • B: Cygnet — system-one-open · Score 73.7 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Cygnet. Intelligence: 72.0 vs 71.1; Calibration: 88.0 vs 87.0; Speed: 83.8 vs 91.0; Cost: 54.7 vs 56.4.50100Intelligence72.0 · 71.1Calibration88.0 · 87.0Speed83.8 · 91.0Cost54.7 · 56.4
0–100, the values in the table. A system with no published axis draws at 0 and says so.

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs Cygnet. Choice · open: 85.7 vs 82.9; Choice · sealed: 87.6 vs 80.3; Noul · open: 47.8 vs 57.9; Noul · sealed: 48.6 vs 55.7; Score · open: 81.2 vs 76.6; Score · sealed: 81.1 vs 73.2.50100Choice · open85.7 · 82.9Choice ·sealed87.6 · 80.3Noul · open47.8 · 57.9Noul · sealed48.6 · 55.7Score · open81.2 · 76.6Score ·sealed81.1 · 73.2
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 904 open and 720 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs Cygnet. Easy: 90.0 vs 89.7; Standard: 80.9 vs 82.4; Judge: 83.5 vs 82.2; Hard: 53.8 vs 58.0.50100Easy90.0 · 89.7Standard80.9 · 82.4Judge83.5 · 82.2Hard53.8 · 58.0
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs Cygnet. Easy: 91.2 vs 88.6; Standard: 79.0 vs 73.5; Judge: 73.7 vs 75.4; Hard: 72.9 vs 65.3.50100Easy91.2 · 88.6Standard79.0 · 73.5Judge73.7 · 75.4Hard72.9 · 65.3
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
All values as a table
SpokeA: Jev 1.13.0B: Cygnet
The four score axes
Intelligence72.071.1
Calibration88.087.0
Speed83.891.0
Cost54.756.4
Competence per request type, open / sealed
Choice · open85.782.9
Choice · sealed87.680.3
Noul · open47.857.9
Noul · sealed48.655.7
Score · open81.276.6
Score · sealed81.173.2
Competence per tier — open set
Easy90.089.7
Standard80.982.4
Judge83.582.2
Hard53.858.0
Competence per tier — sealed set
Easy91.288.6
Standard79.073.5
Judge73.775.4
Hard72.965.3

Axes, request types, latency and cost

Compare:

Every measured system, under the four score axes. Intelligence is 50% open (904 decisions) and 50% sealed (720); Gap = I_open − I_sealed, and the penalty applies only above the field median gap (G_med 5.2) plus 8. The per-type competence, latency and price of the same systems are in Types & cost. On a phone the name column stays put while the table scrolls sideways.

#ASystemScoreIntel.Calib.SpeedCostGapPenalty
1Cygnet73.771.187.091.056.4+2.7×1.000
2Winnow-12B Q873.274.484.186.156.6-3.9×1.000
3Jev 1.13.0API72.172.088.083.854.7-0.9×1.000
4Jev-Omni71.570.582.684.756.1-0.0×1.000
5decider-4b v271.355.885.690.964.5+11.9×1.000
6SemIf68.751.384.090.963.1+5.4×1.000
7spark-s1-4b-v668.262.169.785.560.4+0.7×1.000
8metask-jev-4b67.553.582.789.557.7+6.5×1.000
9Hopper67.549.987.987.262.3-2.3×1.000
10Malkuth-4B66.854.583.286.255.8+7.8×1.000
11reflex 4B65.251.586.868.863.1+7.4×1.000
12jev-local65.256.377.873.058.9+3.8×1.000
13djev (Maisa, diffusion-gemma)64.272.380.491.048.2+0.8×1.000
14Raw Qwen3 4B Instruct 2507 direct logits62.154.152.989.263.4+6.7×1.000
15jqv60.849.186.883.451.2+4.5×1.000
16JevK5 v0.2.058.146.784.990.963.1+6.6×1.000
17Qwen3.5-9B Jev-like data-mix v253.060.480.982.045.7+4.3×1.000
18Standard One 8B47.859.683.292.543.3+0.7×1.000
19NInfer Qwen3.8-Flash-Next mixed47.567.288.588.642.5+6.1×1.000
20swanOne46.671.287.184.742.2-1.5×1.000
21Raw Qwen3 8B direct logits45.251.149.286.845.5+11.3×1.000
22decider-2b45.142.371.594.464.9+12.9×1.000
23system-one44.150.549.490.945.0+6.1×1.000
24system-one-openAPI42.441.671.678.368.3+5.6×1.000
25Autoloops – Gemma 4 31B ITAPI40.576.785.883.939.6+0.4×1.000
26GPT-6 Luna (low reasoning effort)API40.595.394.973.239.1-1.1×1.000
27GPT-6 Luna (default medium reasoning effort)API38.896.295.673.238.3+0.2×1.000
28JevOne38.254.184.990.439.8-0.5×1.000
29kev 4B38.139.967.685.565.8+13.9×0.993
30kev 8B34.248.371.084.140.4+12.0×1.000
31open-alternative-jev33.637.476.991.363.3+10.5×1.000
32Bespoke Nimble 9B31.863.777.283.036.8+2.8×1.000
33Malkuth-2B29.935.575.291.865.6+15.6×0.976
34openjev-sglangAPI29.058.682.978.135.6+5.4×1.000
35decider-35b-a3b27.560.582.091.034.4+7.4×1.000
36local-jev Qwen3.5-4B25.833.782.483.759.3+5.9×1.000
37Open-Jev 9B24.463.881.573.633.1-4.1×1.000
38Decision 2B22.531.386.290.466.4+6.7×1.000
39GPT-5.6 LunaAPI22.494.394.774.130.7-2.3×1.000
40typecastlm21.831.276.792.064.0-15.1×1.000
41JEV Qwen3.5-9B Base NVFP420.132.280.993.547.6+1.3×1.000
42Gemini 3.1 Flash-LiteAPI19.677.674.780.129.8-1.9×1.000
43NInfer Qwen3.8-27B NVFP418.765.585.989.929.1+3.0×1.000
44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.561.186.589.929.1+2.7×1.000
45InstinctAPI18.362.785.081.729.2-1.7×1.000
46OpenJev (thinking, BF16)17.984.283.173.928.5-3.4×1.000
47djev (thinking)17.477.395.772.328.1-1.6×1.000
48LitJev16.358.384.568.228.4-0.5×1.000
49Raw Phi-4 mini direct logits15.227.671.489.253.2+18.3×0.949
50OpenSourceJev13.225.876.072.569.1+4.2×1.000
51reflex-27b13.262.885.969.225.8+0.6×1.000
52Open-Jev 2B9.133.673.675.833.1+10.4×1.000
53GLiNER2 large8.422.542.464.877.6+7.1×1.000
54Qwen3-Reranker-4B7.121.076.379.648.4+17.0×0.962
55DeepSeek V4.1 FlashAPI6.693.796.969.419.1-3.7×1.000
56SimpleJev4.016.746.759.170.3-6.8×1.000
57SimpleJev Qwen3.8-27BAPI3.472.887.274.914.9-1.7×1.000
58decision-machine-1API3.214.881.292.856.3+22.4×0.908
59GLiNER2.5 multi2.714.058.066.686.6+3.6×1.000
60GLiNER22.313.535.670.386.6+7.5×1.000
61JevActAPI1.511.262.576.468.6+16.3×0.969
62CLM-8B1.511.448.593.350.5+11.3×1.000
63kev 0.6B1.310.367.787.280.1+27.0×0.862
64Raw Qwen3 0.6B direct logits1.110.621.590.477.6-1.1×1.000
65GLiNER2.5 small0.99.255.977.486.6+15.5×0.977
66Raw Qwen3 1.7B direct logits0.99.721.690.268.5+5.2×1.000
67Mirror0.25.643.263.689.3+2.3×1.000
68ZeroEntropy zerank-20.14.881.780.348.4+10.4×1.000
69jeff0.14.480.355.881.1+16.3×0.969
70smalljev semantic-v90.13.773.186.560.8+14.5×0.987
71OpenDecision0.02.172.586.679.1+14.7×0.985
72BAAI bge-reranker-v2-m30.00.083.390.559.3+1.7×1.000
73Certo v10.00.087.591.496.9+0.8×1.000
74Decision Fast0.00.076.191.480.1+19.0×0.942
75Alibaba GTE Reranker ModernBERT-base0.00.075.291.569.3+4.5×1.000
76kev 0.5B0.00.065.088.580.1+25.7×0.875
77Laya0.00.073.773.984.9+12.7×1.000
78lev-350m0.00.077.793.680.0+23.5×0.896
79Qwen3.5-0.8B Decision Model0.00.073.671.879.5+7.1×1.000
80Mixedbread mxbai-rerank-base-v20.00.086.689.360.3+3.0×1.000
81Needle 30.00.00.034.161.6+10.0×1.000
82Needle 3, options as tools0.00.00.041.061.6+14.1×0.991
83open-jev-deberta-v3-large0.00.077.168.377.6+5.3×1.000
84Open Jev JSON Canvas0.077.10.085.649.3+4.0×1.000
85openJev Verdict0.00.052.283.986.6+26.7×0.865
86openJev Verdict 1.40.00.080.380.686.6+3.5×1.000
87Qwen3.8 27BAPI0.095.698.156.80.0-5.3×1.000
88verdict-small0.00.059.582.1100.0+2.0×1.000
89Von0.00.083.575.782.7+14.7×0.985
–classifier.devAPIhonorable mention—75.889.282.058.9-2.7×1.000
–JevK5 v0.3v1.5 roster addendum A1—56.388.393.663.1+10.7×1.000
–Plumb-4Bv1.5 roster addendum A1—55.887.493.563.1+10.1×1.000
–Decision 4B v1.2v1.5 roster addendum A1—53.788.693.563.1+5.2×1.000
–Imajev-4Bv1.5 roster addendum A2—53.588.191.163.3+5.5×1.000
–Decision 4B v1.1v1.5 roster addendum A1—53.187.393.563.1+3.1×1.000
–Surogate Rune 26B-A4B v3v1.5 roster addendum A2—69.788.386.049.0-1.6×1.000
–AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1—72.887.787.629.4-1.7×1.000
–AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2—72.886.787.129.4-1.5×1.000
–Eikos-27Bv1.5 roster addendum A1—75.186.387.628.7-1.9×1.000
–SimpleJev Qwen3.6-35B-A3BAPIpartial run—0.078.575.235.1+5.6×1.000
Official order with 95% intervals (89 systems)

JevBench v1.5.0 · headline option A

JevBench Score: 89 ranked systems

Official (A)weighted harmonic mean of four 0–100 axes, Intelligence · Calibration · Speed · Cost = 25 · 25 · 25 · 25, with the low-axis gates · Method ↓

Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).

Whiskers are 95% bootstrap intervals. 75 of 88 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant.

  1. 1Cygnet73.7I 71 · C 87 · S 91 · K 56 · B#2 · C#1 · $0.028
  2. 2Winnow-12B Q873.2I 74 · C 84 · S 86 · K 57 · B#1 · C#2 · $0.028
  3. 3Jev 1.13.0API72.1I 72 · C 88 · S 84 · K 55 · B#3 · C#3 · $0.032
  4. 4Jev-Omni71.5I 70 · C 83 · S 85 · K 56 · B#4 · C#4 · $0.029
  5. 5decider-4b v271.3I 56 · C 86 · S 91 · K 65 · B#5 · C#7 · $0.015
  6. 6SemIf68.7I 51 · C 84 · S 91 · K 63 · B#8 · C#13 · $0.017
  7. 7spark-s1-4b-v668.2I 62 · C 70 · S 86 · K 60 · B#6 · C#5 · $0.021
  8. 8metask-jev-4b67.5I 53 · C 83 · S 89 · K 58 · B#9 · C#10 · $0.026
  9. 9Hopper67.5I 50 · C 88 · S 87 · K 62 · B#12 · C#17 · $0.018
  10. 10Malkuth-4B66.8I 54 · C 83 · S 86 · K 56 · B#10 · C#9 · $0.030
  11. 11reflex 4B65.2I 51 · C 87 · S 69 · K 63 · B#13 · C#14 · $0.017
  12. 12jev-local65.2I 56 · C 78 · S 73 · K 59 · B#11 · C#8 · $0.024
  13. 13djev (Maisa, diffusion-gemma)64.2I 72 · C 80 · S 91 · K 48 · B#7 · C#6 · $0.053
  14. 14Raw Qwen3 4B Instruct 2507 direct logits62.1I 54 · C 53 · S 89 · K 63 · B#14 · C#12 · $0.017
  15. 15jqv60.8I 49 · C 87 · S 83 · K 51 · B#15 · C#19 · $0.042
  16. 16JevK5 v0.2.058.1I 47 · C 85 · S 91 · K 63 · B#16 · C#22 · $0.017
  17. 17Qwen3.5-9B Jev-like data-mix v253.0I 60 · C 81 · S 82 · K 46 · B#17 · C#11 · $0.065
  18. 18Standard One 8B47.8I 60 · C 83 · S 92 · K 43 · B#20 · C#16 · $0.078
  19. 19NInfer Qwen3.8-Flash-Next mixed47.5I 67 · C 89 · S 89 · K 43 · B#18 · C#15 · $0.082
  20. 20swanOne46.6I 71 · C 87 · S 85 · K 42 · B#19 · C#18 · $0.085
  21. 21Raw Qwen3 8B direct logits45.2I 51 · C 49 · S 87 · K 46 · B#21 · C#24 · $0.065
  22. 22decider-2b45.1I 42 · C 72 · S 94 · K 65 · B#26 · C#26 · $0.015
  23. 23system-one44.1I 50 · C 49 · S 91 · K 45 · B#22 · C#27 · $0.068
  24. 24system-one-openAPI42.4I 42 · C 72 · S 78 · K 68 · B#27 · C#29 · $0.011
  25. 25Autoloops – Gemma 4 31B ITAPI40.5I 77 · C 86 · S 84 · K 40 · B#24 · C#20 · tariff$0.103
  26. 26GPT-6 Luna (low reasoning effort)API40.5I 95 · C 95 · S 73 · K 39 · B#23 · C#21 · $0.108
  27. 27GPT-6 Luna (default medium reasoning effort)API38.8I 96 · C 96 · S 73 · K 38 · B#25 · C#23 · $0.114
  28. 28JevOne38.2I 54 · C 85 · S 90 · K 40 · B#28 · C#28 · $0.101
  29. 29kev 4B38.1I 40 · C 68 · S 85 · K 66 · B#29 · C#32 · $0.014
  30. 30kev 8B34.2I 48 · C 71 · S 84 · K 40 · B#30 · C#34 · $0.097
  31. 31open-alternative-jev33.6I 37 · C 77 · S 91 · K 63 · B#32 · C#35 · $0.017
  32. 32Bespoke Nimble 9B31.8I 64 · C 77 · S 83 · K 37 · B#31 · C#25 · $0.128
  33. 33Malkuth-2B29.9I 36 · C 75 · S 92 · K 66 · B#35 · C#37 · $0.014
  34. 34openjev-sglangAPI29.0I 59 · C 83 · S 78 · K 36 · B#33 · C#30 · $0.140
  35. 35decider-35b-a3b27.5I 60 · C 82 · S 91 · K 34 · B#34 · C#31 · $0.154
  36. 36local-jev Qwen3.5-4B25.8I 34 · C 82 · S 84 · K 59 · B#38 · C#43 · $0.023
  37. 37Open-Jev 9B24.4I 64 · C 82 · S 74 · K 33 · B#36 · C#33 · $0.170
  38. 38Decision 2B22.5I 31 · C 86 · S 90 · K 66 · B#41 · C#45 · $0.013
  39. 39GPT-5.6 LunaAPI22.4I 94 · C 95 · S 74 · K 31 · B#37 · C#36 · tariff$0.205
  40. 40typecastlm21.8I 31 · C 77 · S 92 · K 64 · B#45 · C#47 · $0.016
  41. 41JEV Qwen3.5-9B Base NVFP420.1I 32 · C 81 · S 94 · K 48 · B#47 · C#48 · $0.056
  42. 42Gemini 3.1 Flash-LiteAPI19.6I 78 · C 75 · S 80 · K 30 · B#39 · C#38 · tariff$0.219
  43. 43NInfer Qwen3.8-27B NVFP418.7I 65 · C 86 · S 90 · K 29 · B#40 · C#39 · $0.231
  44. 44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.5I 61 · C 86 · S 90 · K 29 · B#43 · C#40 · $0.231
  45. 45InstinctAPI18.3I 63 · C 85 · S 82 · K 29 · B#44 · C#41 · $0.230
  46. 46OpenJev (thinking, BF16)17.9I 84 · C 83 · S 74 · K 29 · B#42 · C#42 · $0.241
  47. 47djev (thinking)17.4I 77 · C 96 · S 72 · K 28 · B#46 · C#44 · $0.249
  48. 48LitJev16.3I 58 · C 84 · S 68 · K 28 · B#48 · C#46 · $0.244
  49. 49Raw Phi-4 mini direct logits15.2I 28 · C 71 · S 89 · K 53 · B#50 · C#50 · $0.036
  50. 50OpenSourceJev13.2I 26 · C 76 · S 73 · K 69 · B#51 · C#51 · $0.011
  51. 51reflex-27b13.2I 63 · C 86 · S 69 · K 26 · B#49 · C#49 · $0.297
  52. 52Open-Jev 2B9.1I 34 · C 74 · S 76 · K 33 · B#52 · C#53 · $0.170
  53. 53GLiNER2 large8.4I 22 · C 42 · S 65 · K 78 · B#54 · C#54 · $0.0056
  54. 54Qwen3-Reranker-4B7.1I 21 · C 76 · S 80 · K 48 · B#55 · C#55 · $0.052
  55. 55DeepSeek V4.1 FlashAPI6.6I 94 · C 97 · S 69 · K 19 · B#53 · C#52 · $0.498
  56. 56SimpleJev4.0I 17 · C 47 · S 59 · K 70 · B#57 · C#57 · $0.0098
  57. 57SimpleJev Qwen3.8-27BAPI3.4I 73 · C 87 · S 75 · K 15 · B#56 · C#56 · $0.687
  58. 58decision-machine-1API3.2I 15 · C 81 · S 93 · K 56 · B#58 · C#58 · $0.029
  59. 59GLiNER2.5 multi2.7I 14 · C 58 · S 67 · K 87 · B#59 · C#59 · $0.0028
  60. 60GLiNER22.3I 13 · C 36 · S 70 · K 87 · B#60 · C#60 · $0.0028
  61. 61JevActAPI1.5I 11 · C 63 · S 76 · K 69 · B#62 · C#61 · $0.011
  62. 62CLM-8B1.5I 11 · C 48 · S 93 · K 51 · B#61 · C#62 · $0.045
  63. 63kev 0.6B1.3I 10 · C 68 · S 87 · K 80 · B#63 · C#63 · $0.0046
  64. 64Raw Qwen3 0.6B direct logits1.1I 11 · C 21 · S 90 · K 78 · B#64 · C#64 · $0.0056
  65. 65GLiNER2.5 small0.9I 9 · C 56 · S 77 · K 87 · B#66 · C#65 · $0.0028
  66. 66Raw Qwen3 1.7B direct logits0.9I 10 · C 22 · S 90 · K 69 · B#65 · C#66 · $0.011
  67. 67Mirror0.2I 6 · C 43 · S 64 · K 89 · B#67 · C#67 · $0.0023
  68. 68ZeroEntropy zerank-20.1I 5 · C 82 · S 80 · K 48 · B#68 · C#68 · $0.052
  69. 69jeff0.1I 4 · C 80 · S 56 · K 81 · B#69 · C#69 · $0.0043
  70. 70smalljev semantic-v90.1I 4 · C 73 · S 86 · K 61 · B#70 · C#70 · $0.020
  71. 71OpenDecision0.0I 2 · C 73 · S 87 · K 79 · B#71 · C#71 · $0.0050
  72. 72BAAI bge-reranker-v2-m30.0I 0 · C 83 · S 91 · K 59 · B#72 · C#72 · $0.023
  73. 73Certo v10.0I 0 · C 88 · S 91 · K 97 · B#73 · C#73 · $0.0013
  74. 74Decision Fast0.0I 0 · C 76 · S 91 · K 80 · B#74 · C#74 · $0.0046
  75. 75Alibaba GTE Reranker ModernBERT-base0.0I 0 · C 75 · S 91 · K 69 · B#75 · C#75 · $0.011
  76. 76kev 0.5B0.0I 0 · C 65 · S 88 · K 80 · B#76 · C#76 · $0.0046
  77. 77Laya0.0I 0 · C 74 · S 74 · K 85 · B#77 · C#77 · $0.0032
  78. 78lev-350m0.0I 0 · C 78 · S 94 · K 80 · B#78 · C#78 · $0.0046
  79. 79Qwen3.5-0.8B Decision Model0.0I 0 · C 74 · S 72 · K 80 · B#79 · C#79 · $0.0048
  80. 80Mixedbread mxbai-rerank-base-v20.0I 0 · C 87 · S 89 · K 60 · B#80 · C#80 · $0.021
  81. 81Needle 30.0I 0 · C 0 · S 34 · K 62 · B#81 · C#81 · $0.019
  82. 82Needle 3, options as tools0.0I 0 · C 0 · S 41 · K 62 · B#82 · C#82 · $0.019
  83. 83open-jev-deberta-v3-large0.0I 0 · C 77 · S 68 · K 78 · B#83 · C#83 · $0.0056
  84. 84Open Jev JSON Canvas0.0I 77 · C 0 · S 86 · K 49 · B#84 · C#84 · $0.049
  85. 85openJev Verdict0.0I 0 · C 52 · S 84 · K 87 · B#85 · C#85 · $0.0028
  86. 86openJev Verdict 1.40.0I 0 · C 80 · S 81 · K 87 · B#86 · C#86 · $0.0028
  87. 87Qwen3.8 27BAPI0.0I 96 · C 98 · S 57 · K 0 · B#87 · C#87 · $2.18
  88. 88verdict-small0.0I 0 · C 59 · S 82 · K 100 · B#88 · C#88 · $0.0009
  89. 89Von0.0I 0 · C 83 · S 76 · K 83 · B#89 · C#89 · $0.0038

Costs are estimates (est.) unless marked tariff.

system-one-openJev rebuildJev (TypeSafe, closed)Raw-logit control (base model)Native-logit decision engineInstruction model, JSON schemaZero-shot classifierReranker (neutral adapter)Closed decision APISmall tool-calling modelService built on JevUnclassified

All three weight options (89 systems)

All three weight options

A is the official headline: equal 25/25/25/25 axis weights and an Intelligence floor of 50. B remains the secondary 40/20/20/20 axis-weight view; C retains equal axes with an Intelligence floor of 60. All three use equal Choice/Noul/Score weights. The CI column is the paired-bootstrap 95% interval of the A score.

#ASystemA · equal (headline)B · 40/20/20/20C · equal, I floor 60#B#CA 95% CI
1Cygnet73.773.273.72172.4–74.5
2Winnow-12B Q873.273.573.21272.0–74.0
3Jev 1.13.0API72.172.172.13371.0–72.6
4Jev-Omni71.571.371.54470.2–72.4
5decider-4b v271.367.561.65769.1–72.3
6SemIf68.764.350.281360.1–70.2
7spark-s1-4b-v668.266.968.26566.3–69.7
8metask-jev-4b67.564.153.691065.4–68.7
9Hopper67.562.946.9121756.3–69.2
10Malkuth-4B66.863.955.010964.9–67.9
11reflex 4B65.261.948.0131458.6–66.3
12jev-local65.263.257.411863.5–66.5
13djev (Maisa, diffusion-gemma)64.264.864.27663.1–65.0
14Raw Qwen3 4B Instruct 2507 direct logits62.160.350.4141259.4–64.2
15jqv60.857.542.2151951.3–64.3
16JevK5 v0.2.058.153.640.4162247.6–68.2
17Qwen3.5-9B Jev-like data-mix v253.052.553.0171151.9–53.9
18Standard One 8B47.847.147.2201646.7–48.5
19NInfer Qwen3.8-Flash-Next mixed47.547.747.5181546.7–47.9
20swanOne46.647.346.6191845.8–46.9
21Raw Qwen3 8B direct logits45.244.632.9212436.9–46.7
22decider-2b45.141.131.3262633.8–54.3
23system-one44.143.431.2222736.8–45.7
24system-one-openAPI42.438.829.5272933.8–51.9
25Autoloops – Gemma 4 31B ITAPI40.541.840.5242040.1–40.8
26GPT-6 Luna (low reasoning effort)API40.543.140.5232140.3–40.7
27GPT-6 Luna (default medium reasoning effort)API38.841.438.8252338.6–38.9
28JevOne38.237.431.1282837.4–38.8
29kev 4B38.134.626.4293227.5–46.8
30kev 8B34.233.123.7303426.8–37.5
31open-alternative-jev33.629.923.3323525.8–41.7
32Bespoke Nimble 9B31.832.331.8312531.2–32.3
33Malkuth-2B29.926.420.8353721.0–39.0
34openjev-sglangAPI29.029.227.7333028.4–29.4
35decider-35b-a3b27.527.727.5343126.9–27.9
36local-jev Qwen3.5-4B25.822.717.9384319.9–33.0
37Open-Jev 9B24.425.024.4363323.9–24.6
38Decision 2B22.519.315.6414516.7–28.8
39GPT-5.6 LunaAPI22.424.222.4373622.3–22.5
40typecastlm21.818.915.2454716.2–28.6
41JEV Qwen3.5-9B Base NVFP420.117.813.9474815.2–25.5
42Gemini 3.1 Flash-LiteAPI19.620.819.6393819.3–19.8
43NInfer Qwen3.8-27B NVFP418.719.318.7403918.5–18.9
44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.518.918.5434018.2–18.6
45InstinctAPI18.318.918.3444118.0–18.5
46OpenJev (thinking, BF16)17.919.317.9424217.8–18.1
47djev (thinking)17.418.517.4464417.3–17.5
48LitJev16.316.715.4484616.0–16.5
49Raw Phi-4 mini direct logits15.213.110.6505010.3–20.7
50OpenSourceJev13.211.29.251519.0–18.3
51reflex-27b13.213.813.2494913.0–13.3
52Open-Jev 2B9.18.46.352536.6–11.5
53GLiNER2 large8.47.25.854545.2–12.3
54Qwen3-Reranker-4B7.15.94.955554.3–10.7
55DeepSeek V4.1 FlashAPI6.67.46.653526.6–6.7
56SimpleJev4.03.32.857572.1–6.7
57SimpleJev Qwen3.8-27BAPI3.43.73.456563.3–3.4
58decision-machine-1API3.22.52.358581.7–5.4
59GLiNER2.5 multi2.72.11.959591.2–5.1
60GLiNER22.31.81.660600.9–4.5
61JevActAPI1.51.11.162610.5–3.2
62CLM-8B1.51.21.061620.5–3.2
63kev 0.6B1.30.90.963630.5–2.5
64Raw Qwen3 0.6B direct logits1.10.90.764640.3–2.6
65GLiNER2.5 small0.90.60.666650.2–2.2
66Raw Qwen3 1.7B direct logits0.90.70.665660.1–2.3
67Mirror0.20.20.267670.0–0.9
68ZeroEntropy zerank-20.10.10.168680.0–0.5
69jeff0.10.10.169690.0–0.5
70smalljev semantic-v90.10.00.070700.0–0.4
71OpenDecision0.00.00.071710.0–0.2
72BAAI bge-reranker-v2-m30.00.00.072720.0–0.0
73Certo v10.00.00.073730.0–0.0
74Decision Fast0.00.00.074740.0–0.0
75Alibaba GTE Reranker ModernBERT-base0.00.00.075750.0–0.0
76kev 0.5B0.00.00.076760.0–0.0
77Laya0.00.00.077770.0–0.0
78lev-350m0.00.00.078780.0–0.0
79Qwen3.5-0.8B Decision Model0.00.00.079790.0–0.0
80Mixedbread mxbai-rerank-base-v20.00.00.080800.0–0.0
81Needle 30.00.00.081810.0–0.0
82Needle 3, options as tools0.00.00.082820.0–0.0
83open-jev-deberta-v3-large0.00.00.083830.0–0.0
84Open Jev JSON Canvas0.00.00.084840.0–0.0
85openJev Verdict0.00.00.085850.0–0.0
86openJev Verdict 1.40.00.00.086860.0–0.0
87Qwen3.8 27BAPI0.00.00.087870.0–0.0
88verdict-small0.00.00.088880.0–0.0
89Von0.00.00.089890.0–0.0

What the run says

  • Among the 56 Jev-class systems, Jev 1.13.0 has the highest Capability, 80.0 (Intelligence 72.0, Calibration 88.0).
  • Cygnet leads the official JevBench v1.5.0 score (option A) with 73.7: Intelligence 71.1, Calibration 87.0, Speed 91.0, Cost 56.4 ($0.028 per 1,000 decisions).
  • The best open or open-planned rebuild, Winnow-12B Q8, is #2 at 73.2 — 0.5 points behind.
  • GPT-6 Luna (default medium reasoning effort) has the highest Intelligence (96.2) but places #27: Speed 73.2, Cost 38.3 — the harmonic mean does not let accuracy buy back a weak axis.
  • The strongest sealed Intelligence is 98.2 (Qwen3.8 27B, #87); sealed items carry half of Intelligence, and an open-minus-sealed gap beyond the field median plus eight points costs Intelligence.
  • 75 of 88 adjacent pairs with published paired-bootstrap comparisons are statistical ties — read the order as a ranking, not the gaps as significant.
  • 9 systems joined by separately hashed roster addenda; they sit outside the frozen v1.5.0 order — their placements against it are in the addendum table below.
  • classifier.dev scores 74.7 but is not ranked: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2).
  • SimpleJev Qwen3.6-35B-A3B did not complete the full suite; they are listed without a rank.

Jev alternatives, open source and self-hosting

The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are Cygnet (#1, 73.7), Winnow-12B Q8 (#2, 73.2), Jev-Omni (#4, 71.5), decider-4b v2 (#5, 71.3). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost — 25/25/25/25 — with an Intelligence floor of 50 and low-axis gates on Speed and Cost. Version v1.5.0 measures 904 open and 720 sealed decisions per system; Choice, Noul and Score each carry a third, sealed items contribute 50% of Intelligence, and an open-minus-sealed gap beyond the field median costs points. Method notes · options B and C.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version or a disclosed roster addendum. For private data, see the custom evaluation options.

What a decision costs

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole typed request — state, rubric and options — not a single token.

How costs are estimated

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured — 3 rows carry a tariff. Systems without one — open weights, author demos, models we ran ourselves — are priced as if a large inference provider hosted them: the list price of the same weights, or the nearest larger sibling or size class when the exact weights are not listed. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

Price rules (v1.5). Only public, bookable list prices that have been in effect for at least 30 days count; a manufacturer's standard, non-promotional launch list price counts from day one, and promotions, subsidies, credits and free tiers never do. The scoring price is never below the market reference price of the system's base model. A system without any eligible price is listed as unpriced — no Cost axis and no score until a price qualifies. A later price change triggers a re-score with a visible note on the row.

  • Cygnet — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Winnow-12B Q8 — ~$0.028 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Jev 1.13.0 — ~$0.032 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • Jev-Omni — ~$0.029 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • decider-4b v2 — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • SemIf — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • spark-s1-4b-v6 — ~$0.021 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • metask-jev-4b — ~$0.026 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Hopper — ~$0.018 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Malkuth-4B — ~$0.030 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • reflex 4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • jev-local — ~$0.024 est. per 1,000 decisions: price floor: base-model reference price applied
  • djev (Maisa, diffusion-gemma) — ~$0.053 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Raw Qwen3 4B Instruct 2507 direct logits — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • jqv — ~$0.042 est. per 1,000 decisions: documented hosted-model estimate
  • JevK5 v0.2.0 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Qwen3.5-9B Jev-like data-mix v2 — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
  • Standard One 8B — ~$0.078 est. per 1,000 decisions: documented hosted-model estimate
  • NInfer Qwen3.8-Flash-Next mixed — ~$0.082 est. per 1,000 decisions: documented hosted-model estimate
  • swanOne — ~$0.085 est. per 1,000 decisions: price floor: base-model reference price applied
  • Raw Qwen3 8B direct logits — ~$0.065 est. per 1,000 decisions: documented hosted-model estimate
  • decider-2b — ~$0.015 est. per 1,000 decisions: documented hosted-model estimate
  • system-one — ~$0.068 est. per 1,000 decisions: documented hosted-model estimate
  • system-one-open — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • GPT-6 Luna (low reasoning effort) — ~$0.108 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • GPT-6 Luna (default medium reasoning effort) — ~$0.114 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • JevOne — ~$0.101 est. per 1,000 decisions: documented hosted-model estimate
  • kev 4B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • kev 8B — ~$0.097 est. per 1,000 decisions: price floor: base-model reference price applied
  • open-alternative-jev — ~$0.017 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Bespoke Nimble 9B — ~$0.128 est. per 1,000 decisions: documented hosted-model estimate
  • Malkuth-2B — ~$0.014 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • openjev-sglang — ~$0.140 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • decider-35b-a3b — ~$0.154 est. per 1,000 decisions: price floor: base-model reference price applied
  • local-jev Qwen3.5-4B — ~$0.023 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Open-Jev 9B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
  • Decision 2B — ~$0.013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • typecastlm — ~$0.016 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • JEV Qwen3.5-9B Base NVFP4 — ~$0.056 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): exact base-model market reference
  • NInfer Qwen3.8-27B NVFP4 — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
  • NInfer Qwen3.8-27B NVFP4 (T=1.5) — ~$0.231 est. per 1,000 decisions: price floor: base-model reference price applied
  • Instinct — ~$0.230 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • OpenJev (thinking, BF16) — ~$0.241 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • djev (thinking) — ~$0.249 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • LitJev — ~$0.244 est. per 1,000 decisions: price floor: base-model reference price applied
  • Raw Phi-4 mini direct logits — ~$0.036 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • OpenSourceJev — ~$0.011 est. per 1,000 decisions: price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • reflex-27b — ~$0.297 est. per 1,000 decisions: price floor: base-model reference price applied
  • Open-Jev 2B — ~$0.170 est. per 1,000 decisions: documented hosted-model estimate
  • GLiNER2 large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Qwen3-Reranker-4B — ~$0.052 est. per 1,000 decisions: ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
  • DeepSeek V4.1 Flash — ~$0.498 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • SimpleJev — ~$0.0098 est. per 1,000 decisions: documented hosted-model estimate
  • SimpleJev Qwen3.8-27B — ~$0.687 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • decision-machine-1 — ~$0.029 est. per 1,000 decisions: operator standard launch list price (interpretation I-1); no exact base-model floor applies
  • GLiNER2.5 multi — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • GLiNER2 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • JevAct — ~$0.011 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • CLM-8B — ~$0.045 est. per 1,000 decisions: price floor: base-model reference price applied
  • kev 0.6B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Raw Qwen3 0.6B direct logits — ~$0.0056 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • GLiNER2.5 small — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Raw Qwen3 1.7B direct logits — ~$0.011 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Mirror — ~$0.0023 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • ZeroEntropy zerank-2 — ~$0.052 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies
  • jeff — ~$0.0043 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • smalljev semantic-v9 — ~$0.020 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • OpenDecision — ~$0.0050 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • BAAI bge-reranker-v2-m3 — ~$0.023 est. per 1,000 decisions: ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3)
  • Certo v1 — ~$0.0013 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Decision Fast — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Alibaba GTE Reranker ModernBERT-base — ~$0.011 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies
  • kev 0.5B — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Laya — ~$0.0032 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • lev-350m — ~$0.0046 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Qwen3.5-0.8B Decision Model — ~$0.0048 est. per 1,000 decisions: documented hosted-model estimate
  • Mixedbread mxbai-rerank-base-v2 — ~$0.021 est. per 1,000 decisions: ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies
  • Needle 3 — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Needle 3, options as tools — ~$0.019 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • open-jev-deberta-v3-large — ~$0.0056 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Open Jev JSON Canvas — ~$0.049 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • openJev Verdict — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • openJev Verdict 1.4 — ~$0.0028 est. per 1,000 decisions: ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies
  • Qwen3.8 27B — ~$2.18 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)
  • verdict-small — ~$0.0009 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • Von — ~$0.0038 est. per 1,000 decisions: documented hosted-model estimate; no exact base-model floor applies
  • classifier.dev — ~$0.023 est. per 1,000 decisions: ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies
  • JevK5 v0.3 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Plumb-4B — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Decision 4B v1.2 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Imajev-4B — ~$0.017 est. per 1,000 decisions: ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B.
  • Decision 4B v1.1 — ~$0.017 est. per 1,000 decisions: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)
  • Surogate Rune 26B-A4B v3 — ~$0.050 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
  • AutoJev-27B (denis-pplx, Qwen3.8-27B) — ~$0.226 est. per 1,000 decisions: documented hosted-model estimate
  • AutoJev-27B (RTX PRO 6000) — ~$0.226 est. per 1,000 decisions: ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.
  • Eikos-27B — ~$0.238 est. per 1,000 decisions: documented hosted-model estimate
  • SimpleJev Qwen3.6-35B-A3B — ~$0.145 est. per 1,000 decisions: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Honorable mentions and Jev wrappers — listed separately, not ranked (1)

Eligibility rule: services that run on Jev itself may be measured and shown as honorable mentions, but are not competitors ranked against Jev and do not enter the field median gap (G_med) or tie markers. classifier.dev (TypeSafe) runs on Jev, so it stays unranked under this rule.

  • classifier.dev (fast tier)APIhonorable mention: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2). Official (A) score 74.7.

Roster addendum: newcomers scored on the same frozen protocol (9)

Added by separately hashed roster addenda before they ran. Same frozen sample, method, price rules and v1.5.0 median gap. These rows stay outside the v1.5.0 order and its tie markers. A and secondary B placement compare each row with the frozen base point estimates only; each row's interval is shown separately and does not establish a tie with a base row or another addendum.

SystemWould place (A)A score · 95% CIWould place (B)B score · 95% CI
JevK5 v0.3v1.5 roster addendum A1#471.9 69.4–72.9#568.1 64.9–69.8
Plumb-4Bv1.5 roster addendum A1#471.6 69.2–72.7#567.7 64.7–69.6
Decision 4B v1.2v1.5 roster addendum A1#670.8 68.5–72.0#766.6 63.7–68.4
Imajev-4Bv1.5 roster addendum A2#670.4 67.8–71.6#766.2 63.1–68.1
Decision 4B v1.1v1.5 roster addendum A1#670.4 66.9–71.6#766.1 62.1–67.9
Surogate Rune 26B-A4B v3v1.5 roster addendum A2#1166.5 65.4–67.0#766.6 65.1–67.5
AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1#4319.5 19.3–19.6#4020.4 20.1–20.6
AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2#4319.5 19.3–19.6#4020.4 20.1–20.6
Eikos-27Bv1.5 roster addendum A1#4418.5 18.3–18.6#4019.5 19.2–19.7

Score followed by its 95% interval. Placements compare point estimates with the frozen base only.

Not ranked: partial, unpriced and unmeasured systems

These systems are part of the 103-system v1.5 roster but have no rank. Their numbers are never shown as zero or free.

Partial runs (1)

  • SimpleJev Qwen3.6-35B-A3BAPIpartial run: Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.

Incomplete or not measured in v1.5 (3)

  • Jobe Qwen3.5-4B (frozen): not measured.
  • mica-v01-4bv1.5 roster addendum A1: not measured.
  • OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16): not measured.

Incomplete and unmeasured systems receive no official rank. Existing results from earlier benchmark versions remain on their frozen version pages.

Method notes: what changed in v1.5

The four axes

Intelligence · 25%
How often answers are right above chance: each type is normalized against its task-specific random baseline (chance = 0, perfect = 100; below-chance tiers can be negative). Choice, Noul and Score count one third each. Easy, Standard, Judge and Hard items count 10%, 20%, 30% and 40%; the open and sealed sets count equally.
Calibration · 25%
How closely stated probabilities match what happens. It uses ECE and TVD for Choice, ECE and Brier for Noul, and normalized RPS plus top-level ECE for Score. The three types count equally; open and sealed items are pooled.
Speed · 25%
Serial response latency on open Standard and Judge items. The p50 and p95 each get a log score: 100 − 20 × log₁₀(seconds ÷ 0.1), then are averaged. Self-hosted and demo endpoints get the published ×2 plus 0.15-second adjustment.
Cost · 25%
Estimated or billed US dollars per 1,000 decisions, using pooled token use across 1,624 decisions and the documented price rules. The log score is 100 − 30 × log₁₀(cost ÷ $0.001). The price reference is $0.001 per 1,000 decisions.

The official score is the weighted harmonic mean of the four axes. Option A gives each axis 25%; option B (40/20/20/20) is a secondary view, while option C keeps equal axes and sets the Intelligence gate at 60. In A and B, Intelligence, Speed and Cost each have a quadratic gate below 50. To limit benchmaxxing on the public items, the open-minus-sealed Intelligence gap may be up to 8 points above the field median (G_med) before a penalty applies. Each further point lowers the multiplier on unpenalized Intelligence by one percentage point.

Frozen method METHOD-v1.5, SHA-256 c25d3d8b8512…; pricing addendum v1.5-M2, SHA-256 2fc44459ef80….

The method owner chose equal axis weights and equal weights for Choice, Noul and Score after reviewing the What-If Lab, preserving continuity with v1.4 and treating the three decision types equally. Disclosed headline amendment: equal-axis, equal-type A, SHA-256 752ddccc4e19…. B remains a secondary view.

  • 1,624 decisions per system: 904 open (601 published) and 720 sealed, drawn fresh from a private pool with the same tier mix as the open set. Sealed counts for 50% of Intelligence: base = 0.5 × I_open + 0.5 × I_sealed.
  • Three request types are scored natively and chance-corrected per item: Choice, Noul and Score each receive one third. Tier weights easy / standard / judge / hard = 10 / 20 / 30 / 40. A type a system does not support is excluded, never scored zero; only full-coverage systems are ranked.
  • Overfit penalty relative to the field: excess = gap − G_med, penalty = max(0, 1 − max(0, excess − 8) / 100). G_med for this batch is 5.2 CC points.
  • Calibration is typed (Choice ECE/TVD, Noul ECE with Brier, Score normalised RPS and top-level ECE), pooled over open and sealed. Speed and Cost formulas are unchanged from v1.4; self-hosted and demo endpoints carry the ×2 + 0.15 s adjustment. A manufacturer's standard, non-promotional launch list price counts from day one, but a newer price cut younger than 30 days does not. Rows without token counts use the measured proxy-token basis. A system without any eligible public, bookable price is listed as unpriced.
  • The frozen 25 Sep DeepInfra snapshot records Qwen3.5-4B as deprecated on 11 Jun 2026 and replaced by Qwen3.5-9B. Its frozen snapshot rates remain the v1.5 M2 reference; price basis tooltips and the correction note disclose this. Pricing disclosure correction SHA-256: 1b660648bd49….
  • Composite: weighted harmonic mean with the Intelligence, Speed and Cost gates below 50 (Intelligence below 60 in option C). The official headline A uses equal 25 / 25 / 25 / 25 axis weights and Intelligence floor 50. B remains the secondary 40 / 20 / 20 / 20 view; C keeps equal axes and Intelligence floor 60. Ties come from the paired bootstrap.
  • Rows marked with a v1.5 roster addendum label were added by separately hashed roster addenda: same frozen sample, method, pricing rules and G_med. They remain outside the base release order and its tie markers.
  • Before every release we review the leaderboard for anomalies and close loopholes with general, documented rules. The page and Git repository provide transparent data and method details; Benchmark Heaven owns its rules.

Data file SHA-256 6b2f6b058b36203c98ec5f585eb8376038bc905f11db944f4e0bcd29c278c643 · scorer output SHA-256 452885de2a84cd5b9ed393d541fd8d6c6a540f9d7f762f9d2df9349383f26173 · run kind official.

Limits
  • 1,624 decisions per system (904 open, 720 sealed) is a measurement, not a census, and it is English-only.
  • The weights are a choice. Option A weights the four axes equally and uses a harmonic mean, so the weakest axis dominates; options B and C are published alternatives and the weight sliders re-score the same axes for exploration — only the official option gives the official score and rank. If a wrong decision costs you more than a slow or expensive one, read the Intelligence column and the per-type competence rather than the score alone.
  • The latency adjustment (×2, +0.15 s on our own servers and demo endpoints) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. Serving under load trades per-user speed for throughput. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo.
  • Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
  • Latency is one origin at one time of day; hosted endpoints, public demos and our own pods are different kinds of latency. Public demo endpoints are shared with everyone else using them.
  • Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version or a disclosed roster addendum rather than silently changing this one.

3D view: three.js r128 (MIT).

Loading capability views…

Loading context-length views…

Previous release: JevBench v1.4.2.2 (frozen results).

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.