JevBench by Benchmark Heaven · released v1.4.2.2

Jev alternatives, compared on the published board

The best Jev alternative depends on your use case. Compare Jev-class decision models using the same released benchmark: Intelligence, Calibration, Speed and Cost, with each row’s evidence and openness notes.

JevBench Scores in the current top five

The Jev row is the reference; the other four rows are current alternatives. The overall score is a composite, so check the separate axes for your use case.

  1. 1Imajev-4B67.4I 52 · C 80 · S 91 · K 60 · ~$0.022 est.
  2. 2Plumb-4B65.8I 53 · C 75 · S 93 · K 56 · ~$0.030 est.
  3. 3decider-4b v264.1I 49 · C 75 · S 93 · K 61 · ~$0.020 est.
  4. 4Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
  5. 5JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
All top-five values as a table
RankSystemScoreIntelligenceCalibrationSpeedCostCost basis
1Imajev-4B67.452.280.490.659.7estimated
2Plumb-4B (crh225, JevK5 v0.2 + LoRA)65.853.075.593.555.8estimated
3decider-4b v2 (Mapika)64.149.475.092.960.9estimated
4Jev 1.13.0 (TypeSafe AI)
Reference system
63.353.176.383.352.0measured
5JevK5 v0.2.062.048.974.591.159.5estimated

Cost evidence is labeled by basis so estimated and announced values are not presented as measured charges.

Looking for an open source Jev alternative?

The published v1.4.2.2 aggregate contains 62 Jev-style systems with public code or weights. “Open” does not imply one shared license or unrestricted commercial use; check the source and exact component terms for your use case. Rerankers, classifiers and general LLM baselines are on the full board.

Highest-ranked open-weight entry: Imajev-4B (rank #1, JevBench Score 67.4). Is Jev itself open source?

The model column uses the published row’s description; where a parameter count was missing, the linked primary model source fills that gap. Recorded hardware is run provenance, not a minimum VRAM guarantee; missing values remain “not stated.”

Showing 62 of 62 published open-weight Jev-class rows.

Open-weight Jev-class systems in JevBench v1.4.2.2
SystemModel and sizePublished resultLicense termsRecorded run setupSource or weights
Imajev-4B

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #1
JevBench Score 67.37
Apache-2.0 (author adapter and server metadata)evaluator-owned Lium GPU; exact model recorded in the run receiptPublished source or weights
Plumb-4B (crh225, JevK5 v0.2 + LoRA)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #2
JevBench Score 65.84
Apache-2.0 (weights and code; NOTICE credits JevK5, Qwen and SemIf); base Qwen3.5-4B Apache-2.0H100 80 GB (evaluator-owned Lium pod)Published source or weights
decider-4b v2 (Mapika)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #3
JevBench Score 64.13
Apache-2.0 (package and weights)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
JevK5 v0.2.0

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #5
JevBench Score 62.04
Apache-2.0 (code/adapter); Apache-2.0 (Qwen3.5-4B base)re-run on a throwaway RunPod pod with the original recipe; deviations in its manifestPublished source or weights
Cygnet (blockbrain, frozen Gemma-4-12B-it)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #6
JevBench Score 61.76
shim MIT; weights Apache-2.0 with Google's Gemma Prohibited Use PolicyRTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
Hopper

Qwen3.5-4B plus HopitAI/hopper LoRA

Rank #7
JevBench Score 59.43
Component-specific terms recorded in RESULT.md; submitted adapter release and Qwen base retain their respective termsour GPU (lium.io RTX A6000 48 GB), local loopback HTTP; serialPublished source or weights
Winnow-12B Q8

google/gemma-4-12B-it LoRA fine-tune, merged and exported as Q8_0 GGUF

Rank #8
JevBench Score 55.58
Apache-2.0, including the applicable Gemma 4 base/derivative licence termsour GPU (lium.io RTX 4090 24 GB), reached over the internet from Germany; serial, one request at a timePublished source or weights
reflex 4B (kshetrajna12)

Qwen/Qwen3.5-4B + kshetrajna12/reflex-qwen3.5-4b-lora

Rank #9
JevBench Score 53.99
MIT (code, adapter); Apache-2.0 (base)our RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from GermanyPublished source or weights
djev (Maisa, diffusion-gemma)

inference method on google/diffusiongemma-26B-A4B-it (one structured denoising read), not a separately trained model

Rank #10
JevBench Score 52.23
Apache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weightsproduction API (api.djev.dev, free preview)Published source or weights
Jev-Omni (akhilaaa3, Gemma-4-12B merged)

google/gemma-4-12B-it fine-tuned and merged, with a trained 256-way decision head

Rank #11
JevBench Score 51.34
Apache-2.0, following Gemma 4; dataset rights stated separately by the authorour RunPod GPU (L40 48 GB, Czechia), reached over the internet from GermanyPublished source or weights
metask-jev-4b

Qwen3.5-4B merged r16 LoRA, candidate-logit readout

Rank #12
JevBench Score 47.78
Apache-2.0our GPU (lium.io RTX 5090 32 GB); serial, in-process candidate-logit inferencePublished source or weights
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)

Qwen/Qwen3.5-4B (frozen, BF16)

Rank #13
JevBench Score 47.69
MIT (code); Qwen3.5 weights Apache-2.0RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1)Published source or weights
local-jev Qwen3.5-4B

Qwen/Qwen3.5-4B text model, zero-shot next-token letter probabilities

Rank #15
JevBench Score 46.80
MIT code; Apache-2.0 Qwen weightsour GPU (lium.io A6000 48 GB), reached over the internet from Germany; serial, one request at a timePublished source or weights
system-one-open (Gemma 4 E2B LoRA on an L4)

google/gemma-4-E2B-it + LoRA

Gemma 4 E2B: 2.3B effective parameters; 5.1B including embeddings. Gemma 4 model card

Rank #16
JevBench Score 45.11
MIT (repository LICENSE; Gemma weights keep Google’s terms)author's public demo endpoint (Modal, L4) — not a production servicePublished source or weights
spark-s1-4b-v6 (Open Spark Jev, abhishek085)

Qwen/Qwen3.5-4B with a LoRA, read at the option-letter logits with a fitted temperature (no extra head)

Rank #17
JevBench Score 44.62
Apache-2.0 (code and weights); base Qwen/Qwen3.5-4B Apache-2.0our RunPod GPU (L40 48 GB, Czechia), reached over the internet from GermanyPublished source or weights
Malkuth-4B (newfull5, Kev post-train)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #18
JevBench Score 44.45
CC-BY-NC-4.0, research use only (XNLI and RACE in the training mix)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
jqv (Qwen3-32B zero-shot)

Qwen/Qwen3-32B, bf16, read as a direct-logit classifier (no fine-tuning)

Rank #19
JevBench Score 44.35
Apache-2.0 (Qwen3-32B weights); serving code publicour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from GermanyPublished source or weights
decider-35b-a3b (Mapika)

Qwen3.5-35B-A3B-Base with a trained decision readout, 34.7B parameters / 3B active

Rank #21
JevBench Score 41.18
Apache-2.0our RunPod GPU (H100 NVL 96 GB), reached over the internetPublished source or weights
OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #23
JevBench Score 40.87
MIT (repository code); Apache-2.0 (Qwen/Qwen3.5-4B base and unsloth/Qwen3.5-4B-GGUF Q4_K_M conversion)re-run on a throwaway RunPod pod with the original recipe; deviations in its manifestPublished source or weights
Malkuth-2B (newfull5, Kev post-train)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #26
JevBench Score 38.93
CC-BY-NC-4.0, research use only (XNLI and RACE in the training mix)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
JEV Qwen3.5-9B Base NVFP4

Qwen3.5-9B Base NVFP4 with compact BF16 decision head

Rank #28
JevBench Score 37.70
Apache-2.0 entrant and upstream checkpointour GPU (lium.io RTX 5090 32 GB), local loopback HTTP, native FP4; serialPublished source or weights
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)

nvidia/diffusiongemma-26B-A4B-it-NVFP4

Rank #29
JevBench Score 36.85
Apache-2.0 (repo and weights)RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1)Published source or weights
kev 4B (research preview)

Qwen3-4B-Base + LoRA + learned pointer head; jaredpalmer/kev-4b

Rank #30
JevBench Score 36.14
Apache-2.0our RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internetPublished source or weights
Decision 2B (FlyMy.AI, v59)

openbmb/MiniCPM5-2B with a trained LoRA adapter and pointer head (26.2M trainable parameters)

Rank #31
JevBench Score 35.80
Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial releaseour RunPod GPU (L40 48 GB, Czechia), reached over the internet from GermanyPublished source or weights
Qwen3.5-9B Jev-like data-mix v2

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #32
JevBench Score 35.24
Apache-2.0 (adapter code and weights; Qwen3.5-9B base is Apache-2.0)re-run on a throwaway RunPod pod with the original recipe; deviations in its manifestPublished source or weights
SimpleJev Qwen3.8-27B

Qwen3.8-27B through SimpleJev's direct-logit classifier

Rank #34
JevBench Score 34.57
Apache-2.0 (Qwen weights); repository licence not statedauthor's public demo endpoint (Featherless Classifier Demo) — not a production servicePublished source or weights
swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4)

Model details not stated in the published row.

Qwen3.8-Flash family: 125B total and 6B active per token; this submission is an NVFP4 variant. QwenCloud model guide

Rank #36
JevBench Score 33.61
patches Apache-2.0, shim MIT; weights under the Qwen licence (LICENSE-NOTICE.md)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
open-alternative-jev (Qwen3.5-4B, IkerMoel)

Qwen/Qwen3.5-4B (frozen, BF16)

Rank #38
JevBench Score 33.24
Apache-2.0 (code and weights)RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1)Published source or weights
jev-local (Qwen3.5-9B)

Qwen/Qwen3.5-9B, frozen, per-option mean log-probability

Rank #39
JevBench Score 32.54
no licence stated in the repository (public code); Apache-2.0 base weightsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from GermanyPublished source or weights
Decision Fast (FlyMy.AI, v53a)

Qwen/Qwen3-0.6B-Base with a trained LoRA adapter and pointer head (10.6M trainable parameters)

Rank #40
JevBench Score 32.49
Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial releaseour RunPod GPU (L40 48 GB, Czechia), reached over the internet from GermanyPublished source or weights
decider-2b (Mapika)

Qwen3.5-2B-Base with a trained decision readout, 1.9B

Rank #41
JevBench Score 30.74
Apache-2.0our RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from GermanyPublished source or weights
jeff (Logan Markewich, GLiFormer 400M)

GLiFormer large (knowledgator/gliformer-large-v1, ~400M) behind a TypeSafe-compatible /v1/systemone server

Rank #42
JevBench Score 30.58
MIT (code); GLiFormer weights per their model cardAMD Ryzen 5 3600 (Sandy), 4 threadsPublished source or weights
Laya (Convai Innovations, ModernBERT-large 421M)

ModernBERT-large encoder + option-marker decision head, 421M, RLCD-trained

Rank #43
JevBench Score 30.25
Apache-2.0AMD Ryzen 5 3600 (Sandy), 4 threadsPublished source or weights
Standard One 8B (Standard Thinking)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #45
JevBench Score 29.07
Apache-2.0 (jev-adapter server, LoRA and merged weights; base Ministral 3 8B Apache-2.0)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
lev-350m (Franck Verrot, LFM2.5-350M)

LiquidAI/LFM2.5-350M with a LoRA and a 6.5M-parameter pointer head (a kev clone)

Rank #46
JevBench Score 28.50
Apache-2.0 (code); weights under LiquidAI's LFM1.0 licence, following the LFM2.5-350M baseour RunPod GPU (L40 48 GB, Czechia), reached over the internet from GermanyPublished source or weights
openjev-sglang (Qwen3.6-35B-A3B on SGLang)

Qwen3.6-35B-A3B

Rank #47
JevBench Score 27.65
no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own termsauthor's public demo endpoint (Modal) — not a production servicePublished source or weights
Von (wfzyx, Option-Marker 395M)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #48
JevBench Score 27.48
Apache-2.0 code; Apache-2.0 weights (answerdotai/ModernBERT-large base)re-run on a throwaway RunPod pod with the original recipe; deviations in its manifestPublished source or weights
kev 8B (research preview)

Qwen3-8B-Base + LoRA + learned pointer head; jaredpalmer/kev-8b

Rank #51
JevBench Score 25.56
Apache-2.0our RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internetPublished source or weights
JevOne

Qwen3.6-35B-A3B BF16 with JevOne bidirectional option-logit mapping

Rank #52
JevBench Score 25.51
Component-specific JevOne/Qwen/SGLang terms recorded in RESULT.mdour GPU (lium.io RTX PRO 6000 Blackwell 96 GB), local loopback HTTP, TP1; serial; independently validated one-card conditionPublished source or weights
typecastlm (Mikhail Gribov, Qwen3.5-4B computed head)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #53
JevBench Score 25.28
Apache-2.0 (package and weights)RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
SimpleJev Qwen3.6-35B-A3B

Qwen3.6-35B-A3B through SimpleJev's direct-logit classifier

Rank #54
JevBench Score 24.86
Apache-2.0 (Qwen weights); repository licence not statedauthor's public demo endpoint (Featherless Classifier Demo) — not a production servicePublished source or weights
kev 0.6B (research preview)

Qwen3-0.6B-Base + LoRA + learned pointer head; jaredpalmer/kev-0.6b

Rank #55
JevBench Score 24.75
Apache-2.0our RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internetPublished source or weights
system-one (Qwen3-8B, Sean Goedecke)

Qwen/Qwen3-8B (frozen, BF16), the model of the author's demos

Rank #57
JevBench Score 23.37
no licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1)Published source or weights
LitJev (Qwen3.8-27B)

Qwen/Qwen3.8-27B, frozen, read at the output head

Rank #59
JevBench Score 19.51
Apache-2.0 (code); Apache-2.0 base weightsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from GermanyPublished source or weights
openJev Verdict 1.4

GLiClass ModernBERT-base fine-tuned decision model, 151M; v1.4 fixed inference engine

Rank #60
JevBench Score 19.00
Apache-2.0AMD Ryzen 5 3600 (Sandy), 4 threadsPublished source or weights
kev 0.5B

Qwen2.5-0.5B + LoRA + learned pointer head; jaredpalmer/kev-0.5b

Rank #61
JevBench Score 18.88
Apache-2.0our RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internetPublished source or weights
Bespoke Nimble 9B (Bespoke Labs)

bespokelabs/Bespoke-Nimble-9B (LoRA, adapter unchanged since 93ec5d6), merged into Qwen/Qwen3.5-9B@c202236 with the author's PEFT safe-merge

Rank #62
JevBench Score 18.66
Apache-2.0 (weights); repository without a licence file as of 19 Sepour RunPod GPU (A40 48 GB, Canada), reached over the internet from GermanyPublished source or weights
openJev Verdict (heman10x, ModernBERT-base 151M)

GLiClass ModernBERT-base (knowledgator/gliclass-modern-base-v2.0) fine-tuned, 151M

Rank #64
JevBench Score 18.09
Apache-2.0AMD Ryzen 5 3600 (Sandy), 4 threadsPublished source or weights
reflex-27b (Qwen3.8-27B)

Qwen3.8-27B, frozen, direct-logit readout averaged across two option orders

Rank #66
JevBench Score 17.84
MIT code; Apache-2.0 Qwen weightsour RunPod GPU (H100 NVL 96 GB), reached over the internetPublished source or weights
djev (thinking)

google/diffusiongemma-26b-a4b-it, BF16; full generation with thinking enabled

Rank #68
JevBench Score 15.20
Apache-2.0our GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a timePublished source or weights
OpenJev (thinking, BF16)

google/diffusiongemma-26b-a4b-it, BF16; OpenJev think=512

Rank #70
JevBench Score 14.83
Apache-2.0our GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a timePublished source or weights
Qwen3.5-0.8B Decision Model (Mourad Ghafiri)

Qwen3.5-0.8B-Base fine-tuned as a JevLite decision model with per-question calibration

Rank #71
JevBench Score 14.54
MIT code and training data; Apache-2.0 model weightsRyzen 5 3600, 4 threadsPublished source or weights
open-jev-deberta-v3-large (local CPU)

microsoft/deberta-v3-large

DeBERTa-v3-large: 304M backbone plus 131M embedding parameters. Microsoft model card

Rank #73
JevBench Score 12.65
Apache-2.0 (model card); DeBERTa-v3 keeps its own termsour CPU (2 threads, Ryzen 5 3600)Published source or weights
smalljev semantic-v9

MiniCPM5-2B-Base, 2.5B dense, with LoRA and native semantic decision heads

Rank #74
JevBench Score 12.31
Apache-2.0our GPU (lium.io A6000 48 GB), reached over the internet from Germany; serial, one request at a timePublished source or weights
Open-Jev 9B (Zefan Cai)

Qwen3.5-9B plus rank-8 LoRA and trained scalar decision head

Rank #77
JevBench Score 11.20
MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projectionour RunPod GPU (H100 80GB HBM3), reached over the internetPublished source or weights
Open-Jev 2B (Zefan Cai)

Qwen3.5-2B plus rank-8 LoRA and trained scalar decision head

Rank #78
JevBench Score 9.98
MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projectionour RunPod GPU (H100 80GB HBM3), reached over the internetPublished source or weights
CLM-8B (Contrastive-LM, clm-latest)

Model details not stated in the published row.

Parameter size not stated in the published row.

Rank #80
JevBench Score 8.57
Apache-2.0 (code and CLM-v0.1-8B head); Qwen3-8B encoder Apache-2.0RTX PRO 6000 (evaluator-owned Lium pod)Published source or weights
SimpleJev (Qwen3.5-0.8B, CPU)

Qwen3.5-0.8B through SimpleJev native assistant-prefill option-logit scorer

Rank #81
JevBench Score 7.46
Apache-2.0 server; Apache-2.0 Qwen3.5-0.8B checkpointSandy shared CPU (4 inference threads)Published source or weights
verdict-small (Manavarya09, multilingual-e5-small 118M)

intfloat/multilingual-e5-small (118M multilingual bi-encoder) fine-tuned on a typed-decision mix; each option is scored against the rendered state by cosine similarity, with temperature scaling and a conformal abstain set on top. One encoder pass per option, no tokens generated.

Rank #84
JevBench Score 5.69
Apache-2.0 (verdictml code and the Manav2op/verdict-small checkpoint over intfloat/multilingual-e5-small)AMD Ryzen 5 3600 (Sandy), 4 threadsPublished source or weights
Mirror

Mirror DeBERTa-v3-large 436M classification-span scorer

Rank #86
JevBench Score 2.11
Apache-2.0 wrapper and Mirror release; upstream DeBERTa/model assets retain their own termshosted submission endpoint, serial from Germany; strict 512-token context rejectionDeBERTa-v3-large base model card

The submitted Mirror head/weights were unavailable at the published repository URL when checked. This link is only the DeBERTa base model, not the complete measured Mirror system.

Certo v1 (AltSlate Labs)

ModernBERT-large with a per-option query/scoring head, ~400M parameters

Rank #90
JevBench Score 0.00
MITour RunPod GPU (GeForce RTX 3090 24 GB, community cloud), reached over the internet from GermanyPublished source or weights
Open Jev JSON Canvas (JoshuaSP)

google/diffusiongemma-26B-A4B-it BF16, one-step JSON canvas

Rank #91
JevBench Score 0.00
MIT code; Apache-2.0 DiffusionGemma weightsour lium.io H100 80 GB; local in-process; serial; seed 0; one denoising stepPublished source or weights

Mirror remains listed for completeness. Its published repository returned 404 when checked, so the table links only to Microsoft’s DeBERTa-v3-large base model card.

See the use-case chooser, including accuracy, speed, cost and self-hosting evidence. · Jev vs Laya, a CPU-measured open model

Frequently asked questions

What does JevBench compare?
JevBench compares published Jev-class decision systems across Intelligence, Calibration, Speed and Cost. This page uses the released v1.4.2.2 aggregate.
Which open-weight Jev alternative scores highest?
Imajev-4B is the highest-ranked open-weight entry in v1.4.2.2: rank 1 overall with a JevBench Score of 67.4. Imajev-4B leads the JevBench Score at 67.37, ahead of Plumb-4B (65.84). The v1.4.2 scoring code and earlier measurement rows are unchanged. The score is a composite, not an accuracy percentage or a guarantee for your workload.
What counts as an open-weight Jev alternative on this page?
All 62 Jev-style rows in v1.4.2.2 whose code or weights are public, with a license note and source link in the published row. The license text is kept per row because code, adapters, base weights and datasets can have different terms, including non-commercial ones.
Does the recorded hardware show a minimum deployment requirement?
No. The table reports the setup used for each benchmark run. 8 of these rows were measured on four CPU threads; most others ran on a rented GPU or the author’s own endpoint. It is not a minimum VRAM or hardware guarantee for another revision, quantization, context length or serving stack.
Are cost and speed values directly measured?
The board labels cost evidence as measured, estimated or announced. Self-hosted rows usually carry an estimated cost from a comparable hosted price, and rows served on our own hardware carry a published speed adjustment (such as ×2 + 0.15 s) that is an assumption, not a measurement.