
Everyday photo
Parcel condition
Question: Is the parcel visibly damaged?
- A. yesCorrect
- B. no
Correct answer: A. yes
Source: Synthetic image created for ImageJevBench · public item promo-parcel-01
Hosting = where inference runs; company = where the provider or lab is registered.
EU hosting means the route's inference runs inside the EU: an EU region, AWS Bedrock's EU cross-region (geo) profiles, Azure's Europe Data Zone, or a provider whose entire public fleet is documented as EU-hosted, each checked per model against the provider's documentation. Global deployments do not count, and neither does an EU billing region, an EU company or an EU control plane on its own. One disclosed company-policy exception stays in, marked “EU equivalent”.
Evidence requirements (benchmark evidence, priced provider, measured task tokens) sit above the table they apply to.
Applies to price views & model offers; benchmark evidence stays unfiltered.
Image benchmark · v0.1.4
A held-out comparison of systems that make decisions from images, from interface targets to everyday scenes.
The frozen benchmark has 684 scored items: 228 public and 456 sealed; 93 further items are retired and not scored. This page shows aggregate sealed results only. It contains no sealed task, image, answer key, or per-item prediction.
Wity-1 under author review. The run used the listed production endpoint, but its response did not identify the deployed build. The author is checking the build; this score may change after a full rerun.
Caveat: The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems.
Whole benchmark, 684 decisions. Bars include all 50 systems. Pink bars are hosted APIs; Wity-1 and Gemma are hosted endpoints without a no-retention claim. A system whose Cost or Calibration axis falls under the gate scores 0; the row says which gate it is.
PQ2_0 + Q8_0 MMProj19.26Ranked by the composite score: equal-weight Intelligence, Calibration, Speed and Cost axes, then the unchanged Jev-class gates. Matched gap is signed public-minus-sealed accuracy within the matched families. No system currently exceeds the 15 pp allowance. Hosted systems are marked API because their providers received sealed images and questions; the Gemma 4 endpoint is identified separately in the exposure note.
| # | System | Composite | Intelligence | Calibration | Speed | Cost | Gap (matched) | Earlier split | Public accuracy | Sealed accuracy | USD / 1,000 | p50 / p95 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Imajev-4B Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132. | 76.39 | 73.77 | 90.52 | 87.59 | 61.18 | -4.2 pp | #1 · 76.39 | 185/228 · 81.1% | 380/456 · 83.3% | USD 0.0197 | 0.099 s / 0.174 s |
| 2 | Wity-1 Under author review: server build ID was not recorded; this score may change after verification. Setting: Wity SystemOne, reasoning=auto API | 74.36 | 83.68 | 87.71 | 76.95 | 57.33 | +2.1 pp | — | 200/228 · 87.7% | 410/456 · 89.9% | USD 0.0264 | 0.710 s / 2.844 s |
| 3 | Jev-Omni | 73.10 | 63.93 | 89.90 | 89.68 | 59.52 | -1.0 pp | #2 · 73.10 | 153/228 · 67.1% | 368/456 · 80.7% | USD 0.0224 | 0.076 s / 0.103 s |
| 4 | NeoHorse Jev 4B | 71.94 | 72.89 | 91.20 | 86.76 | 51.56 | -0.8 pp | #3 · 71.94 | 184/228 · 80.7% | 377/456 · 82.7% | USD 0.0412 | 0.140 s / 0.170 s |
| 5 | Visual-Jev 4B Answer-SFT | 69.81 | 65.43 | 84.13 | 87.52 | 53.48 | -6.2 pp | #4 · 69.81 | 175/228 · 76.8% | 352/456 · 77.2% | USD 0.0355 | 0.101 s / 0.176 s |
| 6 | JPT-4Bkirp / llm2jev | 69.55 | 66.92 | 87.31 | 86.52 | 51.13 | +3.3 pp | #5 · 69.55 | 172/228 · 75.4% | 362/456 · 79.4% | USD 0.0426 | 0.123 s / 0.206 s |
| 7 | imajev 2B | 68.72 | 57.73 | 90.46 | 87.85 | 54.19 | -1.1 pp | #6 · 68.72 | 164/228 · 71.9% | 328/456 · 71.9% | USD 0.0336 | 0.096 s / 0.164 s |
| 8 | Glancefrozen Qwen3-VL-4B | 66.86 | 59.37 | 72.97 | 88.73 | 55.54 | -2.8 pp | #7 · 66.86 | 177/228 · 77.6% | 322/456 · 70.6% | USD 0.0303 | 0.089 s / 0.129 s |
| 9 | AutoJev-27B | 66.85 | 69.33 | 82.43 | 86.79 | 49.37 | +0.1 pp | #8 · 66.85 | 174/228 · 76.3% | 371/456 · 81.4% | USD 0.0487 | 0.141 s / 0.167 s |
| 10 | Mapika decider-2b-vision BF16 | 66.77 | 52.84 | 81.46 | 88.17 | 57.59 | -4.4 pp | #9 · 66.77 | 151/228 · 66.2% | 319/456 · 70.0% | USD 0.0259 | 0.091 s / 0.154 s |
| 11 | jev-spatialFr0zencr4nE | 65.87 | 52.04 | 73.04 | 90.52 | 59.63 | -0.7 pp | #10 · 65.87 | 159/228 · 69.7% | 307/456 · 67.3% | USD 0.0222 | 0.058 s / 0.092 s |
| 12 | JPT-9Bkirp / llm2jev | 65.82 | 75.74 | 91.54 | 85.18 | 48.25 | +1.4 pp | #11 · 65.82 | 187/228 · 82.0% | 387/456 · 84.9% | USD 0.0531 | 0.155 s / 0.255 s |
| 13 | OmniJev 4Btinnel123 | 65.65 | 56.85 | 82.61 | 87.06 | 50.65 | -3.0 pp | #12 · 65.65 | 163/228 · 71.5% | 325/456 · 71.3% | USD 0.0442 | 0.131 s / 0.164 s |
| 14 | Reflex 4Breleased stable configuration | 65.38 | 63.34 | 83.68 | 86.14 | 49.35 | -0.6 pp | #13 · 65.38 | 184/228 · 80.7% | 333/456 · 73.0% | USD 0.0488 | 0.154 s / 0.191 s |
| 15 | OmniJev-Qwen3.5-9B-v4tzcfly | 65.14 | 52.17 | 69.59 | 90.34 | 59.54 | +0.1 pp | #14 · 65.14 | 149/228 · 65.4% | 318/456 · 69.7% | USD 0.0223 | 0.063 s / 0.093 s |
| 16 | Jevify Gemma 4 26B-A4B | 65.00 | 71.64 | 87.97 | 86.24 | 48.37 | -1.2 pp | #15 · 65.00 | 189/228 · 82.9% | 366/456 · 80.3% | USD 0.0526 | 0.156 s / 0.182 s |
| 17 | Surogate Rune 26B-A4B v3 | 63.82 | 74.91 | 87.77 | 85.93 | 47.81 | +0.9 pp | #16 · 63.82 | 191/228 · 83.8% | 379/456 · 83.1% | USD 0.0549 | 0.163 s / 0.193 s |
| 18 | JevAny-27B SFT | 63.79 | 69.57 | 90.02 | 86.11 | 48.05 | +4.7 pp | #17 · 63.79 | 177/228 · 77.6% | 369/456 · 80.9% | USD 0.0539 | 0.161 s / 0.185 s |
| 19 | JevAny-27B RLCR | 63.46 | 70.23 | 89.99 | 86.06 | 47.90 | +4.7 pp | #18 · 63.46 | 178/228 · 78.1% | 371/456 · 81.4% | USD 0.0545 | 0.162 s / 0.186 s |
| 20 | Jev-Vision 8BSeanLiu | 63.39 | 64.18 | 63.14 | 85.86 | 49.99 | -1.0 pp | #19 · 63.39 | 181/228 · 79.4% | 340/456 · 74.6% | USD 0.0465 | 0.149 s / 0.215 s |
| 21 | shisa-de-1 | 63.20 | 72.44 | 91.82 | 85.85 | 47.60 | +1.7 pp | #20 · 63.20 | 182/228 · 79.8% | 377/456 · 82.7% | USD 0.0558 | 0.165 s / 0.195 s |
| 22 | OmniJev 2Btinnel123 | 62.84 | 49.47 | 80.53 | 88.53 | 54.40 | -3.9 pp | #21 · 62.84 | 163/228 · 71.5% | 291/456 · 63.8% | USD 0.0331 | 0.097 s / 0.130 s |
| 23 | CUA-S1 4B | 62.63 | 60.17 | 88.57 | 85.39 | 48.55 | +0.5 pp | #22 · 62.63 | 170/228 · 74.6% | 333/456 · 73.0% | USD 0.0519 | 0.150 s / 0.246 s |
| 24 | Standard One 8B | 59.80 | 52.77 | 92.06 | 84.91 | 48.26 | -1.6 pp | #23 · 59.80 | 168/228 · 73.7% | 301/456 · 66.0% | USD 0.0530 | 0.142 s / 0.298 s |
| 25 | Jevify Qwen3-VL-2B T2 | 59.02 | 47.31 | 86.83 | 89.52 | 59.38 | +1.6 pp | #24 · 59.02 | 164/228 · 71.9% | 280/456 · 61.4% | USD 0.0226 | 0.062 s / 0.129 s |
| 26 | imajev 9B | 58.33 | 74.75 | 89.27 | 84.13 | 46.06 | +1.3 pp | #25 · 58.33 | 197/228 · 86.4% | 372/456 · 81.6% | USD 0.0628 | 0.188 s / 0.292 s |
| 27 | Visual-Jev genericzero-shot Qwen3.5-4B baseline | 57.73 | 60.24 | 88.09 | 83.98 | 46.97 | +0.3 pp | #26 · 57.73 | 178/228 · 78.1% | 325/456 · 71.3% | USD 0.0586 | 0.179 s / 0.319 s |
| 28 | djev-distill-v4tarsur385 | 50.04 | 52.54 | 84.81 | 82.98 | 45.11 | +7.9 pp | #27 · 50.04 | 142/228 · 62.3% | 327/456 · 71.7% | USD 0.0676 | 0.195 s / 0.392 s |
| 29 | djev-spark NVFP4 | 49.38 | 49.10 | 84.30 | 83.51 | 45.95 | +3.4 pp | #28 · 49.38 | 146/228 · 64.0% | 307/456 · 67.3% | USD 0.0633 | 0.173 s / 0.374 s |
| 30 | diffusiongemma-26b djev v10 step160snowicarus | 48.76 | 57.19 | 74.30 | 82.80 | 44.65 | +2.6 pp | #29 · 48.76 | 152/228 · 66.7% | 338/456 · 74.1% | USD 0.0700 | 0.203 s / 0.397 s |
| 31 | vjev-visionyah01 | 47.67 | 43.18 | 83.80 | 90.62 | 60.78 | -2.2 pp | #30 · 47.67 | 140/228 · 61.4% | 286/456 · 62.7% | USD 0.0203 | 0.058 s / 0.088 s |
| 32 | Autoloops – Gemma 4 31B ITAPI | 47.14 | 79.27 | 82.38 | 73.89 | 42.65 | +0.7 pp | #31 · 47.14 | 193/228 · 84.6% | 397/456 · 87.1% | USD 0.0816 | 1.425 s / 2.868 s |
| 33 | djev-dev BF16 | 46.64 | 49.60 | 78.02 | 83.00 | 44.69 | +6.6 pp | #32 · 46.64 | 153/228 · 67.1% | 302/456 · 66.2% | USD 0.0698 | 0.193 s / 0.392 s |
| 34 | Winnow-12BQ8_0 + F16 vision projector | 43.56 | 67.40 | 86.23 | 82.12 | 41.35 | +2.9 pp | #33 · 43.56 | 177/228 · 77.6% | 359/456 · 78.7% | USD 0.0902 | 0.268 s / 0.372 s |
| 35 | JPT-0.8Bkirp / llm2jev | 30.51 | 36.27 | 78.06 | 89.14 | 57.50 | -0.3 pp | #34 · 30.51 | 120/228 · 52.6% | 275/456 · 60.3% | USD 0.0261 | 0.074 s / 0.129 s |
| 36 | Jevify Gemma 4 E4B | 29.76 | 35.70 | 76.68 | 90.67 | 60.86 | -1.7 pp | #35 · 29.76 | 129/228 · 56.6% | 263/456 · 57.7% | USD 0.0202 | 0.057 s / 0.088 s |
| 37 | Standard One 3B | 29.01 | 35.76 | 83.90 | 86.52 | 52.36 | -1.1 pp | #36 · 29.01 | 136/228 · 59.6% | 256/456 · 56.1% | USD 0.0387 | 0.100 s / 0.244 s |
| 38 | Gevva E4Btext checkpoint, image input | 26.83 | 35.94 | 69.91 | 85.82 | 49.06 | +1.4 pp | #37 · 26.83 | 132/228 · 57.9% | 261/456 · 57.2% | USD 0.0499 | 0.158 s / 0.206 s |
| 39 | OmniJev 0.8Btinnel123 | 26.81 | 34.57 | 76.47 | 88.95 | 55.39 | +4.4 pp | #38 · 26.81 | 148/228 · 64.9% | 238/456 · 52.2% | USD 0.0307 | 0.088 s / 0.120 s |
| 40 | Gevva E2B multimodal | 21.85 | 32.10 | 75.97 | 86.77 | 51.03 | -3.5 pp | #39 · 21.85 | 115/228 · 50.4% | 261/456 · 57.2% | USD 0.0429 | 0.132 s / 0.179 s |
| 41 | JEVisiondivyanshx11, visual route | 21.50 | 35.43 | 44.40 | 85.45 | 47.30 | +2.3 pp | #40 · 21.50 | 123/228 · 53.9% | 268/456 · 58.8% | USD 0.0571 | 0.170 s / 0.216 s |
| 42 | Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj | 19.26 | 74.91 | 90.71 | 73.62 | 29.43 | +1.9 pp | #41 · 19.26 | 191/228 · 83.8% | 379/456 · 83.1% | USD 0.2252 | 0.799 s / 1.167 s |
| 43 | Qevi-2BMeerDevelopment | 17.44 | 29.54 | 55.67 | 89.37 | 58.64 | +0.1 pp | #42 · 17.44 | 144/228 · 63.2% | 219/456 · 48.0% | USD 0.0239 | 0.068 s / 0.127 s |
| 44 | GPT-6 Lunalow reasoning effort Setting: OpenRouter reasoning.effort=low Cost receipts cover 91.8% of calls API | 16.06 | 81.10 | 88.31 | 65.70 | 27.48 | +12.0 pp | #43 · 16.06 | 203/228 · 89.0% | 395/456 · 86.6% | USD 0.2166 | 2.905 s / 9.256 s |
| 45 | GPT-5.6 Luna Cost receipts cover 99.6% of calls API | 11.05 | 91.81 | 95.78 | 62.58 | 23.49 | -0.8 pp | #44 · 11.05 | 211/228 · 92.5% | 436/456 · 95.6% | USD 0.3524 | 2.878 s / 19.176 s |
| 46 | Gemini 3.1 Flash LiteAPI | 9.58 | 84.28 | 91.44 | 64.49 | 22.31 | -0.2 pp | #45 · 9.58 | 194/228 · 85.1% | 419/456 · 91.9% | USD 0.3888 | 1.799 s / 19.775 s |
| 47 | PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor) | 0.11 | 4.42 | 44.47 | 89.69 | 59.35 | -1.0 pp | #46 · 0.11 | 80/228 · 35.1% | 170/456 · 37.3% | USD 0.0226 | 0.067 s / 0.115 s |
| 48 | Gemini 3.8 Flash0.0 · gated by Cost (USD 2.06 per 1,000 decisions)API | 0.00 | 88.47 | 72.98 | 58.78 | 0.58 | +0.8 pp | #47 · 0.00 | 202/228 · 88.6% | 430/456 · 94.3% | USD 2.0599 | 4.832 s / 27.423 s |
| 49 | OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported) | 0.00 | 63.58 | 0.00 | 79.54 | 40.89 | +3.9 pp | #48 · 0.00 | 187/228 · 82.0% | 331/456 · 72.6% | USD 0.0934 | 0.317 s / 0.634 s |
| 50 | OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported) | 0.00 | 67.21 | 0.00 | 79.49 | 41.09 | +0.4 pp | #49 · 0.00 | 180/228 · 78.9% | 355/456 · 77.9% | USD 0.0920 | 0.321 s / 0.635 s |
Search any two measured systems. The radar uses the same four 0–100 score axes as the ranking; further out is better.
| Axis | Imajev-4B | Wity-1 |
|---|---|---|
| Intelligence | 73.8 | 83.7 |
| Calibration | 90.5 | 87.7 |
| Speed | 87.6 | 76.9 |
| Cost | 61.2 | 57.3 |
These eight examples are from the public split. Each card shows the image, question, options and correct answer; licensed sources are credited below their image.

Everyday photo
Question: Is the parcel visibly damaged?
Correct answer: A. yes
Source: Synthetic image created for ImageJevBench · public item promo-parcel-01

Everyday photo
Question: Is the receipt readable?
Correct answer: A. yes
Source: Synthetic image created for ImageJevBench · public item promo-receipt-00

Computer use · spreadsheet
Question: Goal: change selected cells to type “Text”. Which labelled marker should be clicked?
Correct answer: B. Click marker B
Source: ScreenSpot-Pro · excel_macos_0 · MIT · 210e78d38442 · public item mm-195 · Licence text: MITAdapted: five labelled click markers were drawn on the source screenshot.

Browser
Question: Goal: open a new tab. Which labelled marker should be clicked?
Correct answer: E. Click marker E
Source: ScreenSpot · item 15 · Apache-2.0 · 0be08781e2e1 · public item mm-135 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

Mobile app · translation
Question: Goal: translate. Which labelled marker should be clicked?
Correct answer: D. Click marker D
Source: ScreenSpot · item 341 · Apache-2.0 · 0be08781e2e1 · public item mm-118 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

Computer use · file manager
Question: Goal: copy the file. Which labelled marker should be clicked?
Correct answer: A. Click marker A
Source: ScreenSpot · item 19 · Apache-2.0 · 0be08781e2e1 · public item mm-139 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

Geometry
Question: Find the perimeter of the parallelogram.
Correct answer: A. 78
Source: Geometry3K · source item 8 (public item mm-029) · MIT · fd21e533e1e5 · public item mm-029 · Licence text: MIT

FinQA · financial table
Question: What is the net change in net revenue during 2015 for Entergy Corporation?
Correct answer: C. 94
Source: FinQA · table item · MIT annotations; CDLA-Permissive-1.0 table data · 3d6a736bc67e · public item mm-061 · Licence text: MIT, CDLA-Permissive-1.0Table data: IBM FinTabNet (CDLA-Permissive-1.0). The table was re-rendered as an image by ImageJevBench; no original filing page is shown.
Screenshot images are adapted from the credited datasets (labelled markers added); the FinQA table is re-rendered by ImageJevBench. Licences: Apache-2.0 (ScreenSpot), MIT (ScreenSpot-Pro, Geometry3K, FinQA annotations), CDLA-Permissive-1.0 (FinTabNet table data). App and website content shown in screenshots belongs to its respective owners. No Mind2Web or Android-in-the-Wild image is shown.
Each track is ranked on its own public and sealed items. The licensed core and synthetic everyday-photo results remain separately visible.
139 public · 361 sealed: 62 real-source items and 299 fresh synthetic pool items (documents, charts, inventory, safety).
| # | System | Composite | Intelligence | Calibration | Speed | Cost | Public accuracy | Sealed accuracy | USD / 1,000 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Imajev-4B Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132. | 75.49 | 70.23 | 86.03 | 88.02 | 63.44 | 104/139 · 74.8% | 290/361 · 80.3% | USD 0.0165 |
| 2 | Wity-1 Under author review: server build ID was not recorded; this score may change after verification. Setting: Wity SystemOne, reasoning=auto API | 72.81 | 81.53 | 84.94 | 76.45 | 56.11 | 114/139 · 82.0% | 321/361 · 88.9% | USD 0.0290 |
| 3 | NeoHorse Jev 4B | 70.85 | 69.64 | 86.15 | 87.24 | 52.54 | 103/139 · 74.1% | 289/361 · 80.1% | USD 0.0382 |
| 4 | JPT-9Bkirp / llm2jev | 70.05 | 69.92 | 88.11 | 85.74 | 50.54 | 99/139 · 71.2% | 295/361 · 81.7% | USD 0.0445 |
| 5 | Visual-Jev 4B Answer-SFT | 69.75 | 62.10 | 84.09 | 87.96 | 55.60 | 99/139 · 71.2% | 265/361 · 73.4% | USD 0.0302 |
| 6 | Jev-Omni | 69.71 | 55.96 | 87.32 | 89.67 | 59.17 | 71/139 · 51.1% | 276/361 · 76.5% | USD 0.0230 |
| 7 | JPT-4Bkirp / llm2jev | 67.96 | 60.32 | 83.09 | 86.94 | 53.34 | 87/139 · 62.6% | 273/361 · 75.6% | USD 0.0359 |
| 8 | imajev 2B | 66.70 | 51.86 | 86.58 | 88.24 | 56.17 | 85/139 · 61.2% | 243/361 · 67.3% | USD 0.0289 |
| 9 | Reflex 4Breleased stable configuration | 65.75 | 59.22 | 82.52 | 86.50 | 49.90 | 103/139 · 74.1% | 249/361 · 69.0% | USD 0.0468 |
| 10 | Glancefrozen Qwen3-VL-4B | 65.04 | 55.32 | 70.43 | 88.63 | 55.74 | 99/139 · 71.2% | 239/361 · 66.2% | USD 0.0299 |
| 11 | CUA-S1 4B | 64.27 | 52.91 | 83.32 | 85.89 | 50.79 | 91/139 · 65.5% | 245/361 · 67.9% | USD 0.0437 |
| 12 | AutoJev-27B | 63.96 | 62.27 | 79.80 | 86.64 | 49.17 | 89/139 · 64.0% | 278/361 · 77.0% | USD 0.0495 |
| 13 | imajev 9B | 63.96 | 69.87 | 85.89 | 84.79 | 48.33 | 111/139 · 79.9% | 280/361 · 77.6% | USD 0.0528 |
| 14 | Jevify Gemma 4 26B-A4B | 63.21 | 66.90 | 84.27 | 86.23 | 48.32 | 105/139 · 75.5% | 276/361 · 76.5% | USD 0.0528 |
| 15 | OmniJev 4Btinnel123 | 63.07 | 50.49 | 80.62 | 87.05 | 50.69 | 84/139 · 60.4% | 239/361 · 66.2% | USD 0.0440 |
| 16 | Surogate Rune 26B-A4B v3 | 62.29 | 70.93 | 83.81 | 85.90 | 47.77 | 107/139 · 77.0% | 289/361 · 80.1% | USD 0.0551 |
| 17 | Standard One 8B | 61.90 | 49.60 | 83.84 | 85.52 | 50.45 | 95/139 · 68.3% | 222/361 · 61.5% | USD 0.0448 |
| 18 | shisa-de-1 | 61.88 | 68.36 | 90.46 | 85.83 | 47.52 | 99/139 · 71.2% | 289/361 · 80.1% | USD 0.0562 |
| 19 | JevAny-27B SFT | 60.26 | 56.93 | 88.76 | 86.10 | 48.03 | 90/139 · 64.7% | 276/361 · 76.5% | USD 0.0540 |
| 20 | JevAny-27B RLCR | 59.97 | 57.71 | 88.12 | 86.05 | 47.89 | 91/139 · 65.5% | 278/361 · 77.0% | USD 0.0546 |
| 21 | Jev-Vision 8BSeanLiu | 59.86 | 57.81 | 53.93 | 86.47 | 51.49 | 97/139 · 69.8% | 251/361 · 69.5% | USD 0.0414 |
| 22 | Visual-Jev genericzero-shot Qwen3.5-4B baseline | 59.10 | 55.06 | 81.45 | 84.56 | 48.24 | 99/139 · 71.2% | 238/361 · 65.9% | USD 0.0531 |
| 23 | Mapika decider-2b-vision BF16 | 54.83 | 46.29 | 78.59 | 88.57 | 59.09 | 75/139 · 54.0% | 234/361 · 64.8% | USD 0.0231 |
| 24 | OmniJev 2Btinnel123 | 52.86 | 46.03 | 78.46 | 88.50 | 54.44 | 92/139 · 66.2% | 212/361 · 58.7% | USD 0.0330 |
| 25 | diffusiongemma-26b djev v10 step160snowicarus | 46.15 | 50.22 | 67.68 | 82.80 | 44.66 | 71/139 · 51.1% | 254/361 · 70.4% | USD 0.0699 |
| 26 | Autoloops – Gemma 4 31B ITAPI | 45.99 | 75.10 | 76.49 | 75.04 | 42.62 | 107/139 · 77.0% | 305/361 · 84.5% | USD 0.0818 |
| 27 | Jevify Qwen3-VL-2B T2 | 44.25 | 41.87 | 82.83 | 89.81 | 61.27 | 88/139 · 63.3% | 201/361 · 55.7% | USD 0.0195 |
| 28 | Winnow-12BQ8_0 + F16 vision projector | 42.83 | 59.40 | 80.33 | 82.49 | 41.81 | 89/139 · 64.0% | 267/361 · 74.0% | USD 0.0870 |
| 29 | jev-spatialFr0zencr4nE | 41.74 | 41.80 | 66.00 | 90.49 | 59.32 | 74/139 · 53.2% | 218/361 · 60.4% | USD 0.0227 |
| 30 | OmniJev-Qwen3.5-9B-v4tzcfly | 38.97 | 40.93 | 60.73 | 90.34 | 59.48 | 64/139 · 46.0% | 227/361 · 62.9% | USD 0.0224 |
| 31 | djev-distill-v4tarsur385 | 36.03 | 43.88 | 78.32 | 83.00 | 45.16 | 60/139 · 43.2% | 245/361 · 67.9% | USD 0.0673 |
| 32 | djev-spark NVFP4 | 32.16 | 41.11 | 80.26 | 83.55 | 45.81 | 67/139 · 48.2% | 224/361 · 62.0% | USD 0.0640 |
| 33 | vjev-visionyah01 | 30.79 | 36.10 | 79.43 | 90.70 | 60.97 | 66/139 · 47.5% | 206/361 · 57.1% | USD 0.0200 |
| 34 | djev-dev BF16 | 29.98 | 41.36 | 72.41 | 83.04 | 44.55 | 71/139 · 51.1% | 220/361 · 60.9% | USD 0.0705 |
| 35 | Standard One 3B | 24.68 | 33.34 | 81.21 | 86.99 | 54.85 | 72/139 · 51.8% | 188/361 · 52.1% | USD 0.0320 |
| 36 | Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj | 18.95 | 68.86 | 87.67 | 74.16 | 29.48 | 103/139 · 74.1% | 286/361 · 79.2% | USD 0.2243 |
| 37 | JPT-0.8Bkirp / llm2jev | 18.24 | 29.32 | 76.00 | 89.37 | 58.85 | 49/139 · 35.3% | 201/361 · 55.7% | USD 0.0235 |
| 38 | Jevify Gemma 4 E4B | 17.74 | 28.89 | 75.52 | 90.71 | 61.02 | 59/139 · 42.4% | 187/361 · 51.8% | USD 0.0199 |
| 39 | GPT-6 Lunalow reasoning effort Setting: OpenRouter reasoning.effort=low Cost receipts cover 90.6% of calls API | 17.10 | 76.62 | 84.82 | 66.92 | 28.33 | 114/139 · 82.0% | 310/361 · 85.9% | USD 0.1955 |
| 40 | OmniJev 0.8Btinnel123 | 16.67 | 28.51 | 73.42 | 88.96 | 55.38 | 74/139 · 53.2% | 167/361 · 46.3% | USD 0.0307 |
| 41 | Gevva E2B multimodal | 15.93 | 28.26 | 76.26 | 86.47 | 49.95 | 53/139 · 38.1% | 192/361 · 53.2% | USD 0.0466 |
| 42 | Gevva E4Btext checkpoint, image input | 15.68 | 29.27 | 72.28 | 85.58 | 47.97 | 61/139 · 43.9% | 186/361 · 51.5% | USD 0.0542 |
| 43 | Qevi-2BMeerDevelopment | 15.07 | 27.75 | 54.95 | 89.80 | 60.87 | 83/139 · 59.7% | 153/361 · 42.4% | USD 0.0201 |
| 44 | JEVisiondivyanshx11, visual route | 11.89 | 27.82 | 37.25 | 85.60 | 47.85 | 50/139 · 36.0% | 194/361 · 53.7% | USD 0.0548 |
| 45 | GPT-5.6 Luna Cost receipts cover 99.6% of calls API | 10.90 | 89.58 | 94.00 | 63.08 | 23.40 | 122/139 · 87.8% | 342/361 · 94.7% | USD 0.3550 |
| 46 | Gemini 3.1 Flash LiteAPI | 9.06 | 79.41 | 85.56 | 65.17 | 21.96 | 105/139 · 75.5% | 324/361 · 89.8% | USD 0.3994 |
| 47 | PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor) | 0.04 | 3.00 | 48.35 | 90.12 | 60.99 | 32/139 · 23.0% | 121/361 · 33.5% | USD 0.0200 |
| 48 | Gemini 3.8 Flash0.0 · gated by Cost (USD 2.24 per 1,000 decisions)API | 0.00 | 85.11 | 73.67 | 59.41 | 0.00 | 113/139 · 81.3% | 336/361 · 93.1% | USD 2.2351 |
| 49 | OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported) | 0.00 | 60.06 | 0.00 | 80.20 | 41.25 | 104/139 · 74.8% | 251/361 · 69.5% | USD 0.0909 |
| 50 | OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported) | 0.00 | 61.39 | 0.00 | 80.57 | 41.51 | 96/139 · 69.1% | 266/361 · 73.7% | USD 0.0890 |
89 public images from the promo set; 95 sealed: 61 variants from the reviewed v0.1 candidate pool across 17 matched situations and 34 fresh pool photos. Ambiguous labels were dropped after visual, two-model and gold-blind human checks. No brands and no focused faces.
| # | System | Composite | Intelligence | Calibration | Speed | Cost | Public accuracy | Sealed accuracy | USD / 1,000 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | OmniJev-Qwen3.5-9B-v4tzcfly | 80.48 | 91.76 | 90.96 | 90.40 | 59.71 | 85/89 · 95.5% | 91/95 · 95.8% | USD 0.0220 |
| 2 | jev-spatialFr0zencr4nE | 80.25 | 89.16 | 90.38 | 90.57 | 60.52 | 85/89 · 95.5% | 89/95 · 93.7% | USD 0.0207 |
| 3 | Jev-Omni | 79.86 | 90.77 | 87.74 | 89.74 | 60.50 | 82/89 · 92.1% | 92/95 · 96.8% | USD 0.0207 |
| 4 | Wity-1 Under author review: server build ID was not recorded; this score may change after verification. Setting: Wity SystemOne, reasoning=auto API | 78.44 | 89.92 | 90.13 | 80.10 | 61.38 | 86/89 · 96.6% | 89/95 · 93.7% | USD 0.0194 |
| 5 | Imajev-4B Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132. | 77.90 | 87.41 | 93.84 | 86.56 | 56.50 | 81/89 · 91.0% | 90/95 · 94.7% | USD 0.0282 |
| 6 | AutoJev-27B | 74.24 | 94.36 | 86.21 | 86.90 | 49.93 | 85/89 · 95.5% | 93/95 · 97.9% | USD 0.0467 |
| 7 | vjev-visionyah01 | 73.38 | 69.09 | 80.60 | 90.55 | 60.28 | 74/89 · 83.1% | 80/95 · 84.2% | USD 0.0211 |
| 8 | Mapika decider-2b-vision BF16 | 73.23 | 77.11 | 85.15 | 87.25 | 54.20 | 76/89 · 85.4% | 85/95 · 89.5% | USD 0.0336 |
| 9 | Jevify Qwen3-VL-2B T2 | 72.88 | 69.31 | 89.98 | 88.81 | 55.30 | 76/89 · 85.4% | 79/95 · 83.2% | USD 0.0309 |
| 10 | imajev 2B | 72.67 | 79.39 | 91.79 | 86.91 | 49.98 | 79/89 · 88.8% | 85/95 · 89.5% | USD 0.0465 |
| 11 | Glancefrozen Qwen3-VL-4B | 72.25 | 76.03 | 78.33 | 88.71 | 55.02 | 78/89 · 87.6% | 83/95 · 87.4% | USD 0.0316 |
| 12 | OmniJev 4Btinnel123 | 71.83 | 80.69 | 83.24 | 87.10 | 50.51 | 79/89 · 88.8% | 86/95 · 90.5% | USD 0.0446 |
| 13 | NeoHorse Jev 4B | 71.13 | 84.81 | 92.32 | 86.65 | 49.21 | 81/89 · 91.0% | 88/95 · 92.6% | USD 0.0493 |
| 14 | Jevify Gemma 4 E4B | 69.29 | 60.84 | 72.89 | 90.62 | 60.46 | 70/89 · 78.7% | 76/95 · 80.0% | USD 0.0208 |
| 15 | Jevify Gemma 4 26B-A4B | 69.13 | 89.70 | 90.20 | 86.28 | 48.50 | 84/89 · 94.4% | 90/95 · 94.7% | USD 0.0521 |
| 16 | OmniJev 2Btinnel123 | 68.85 | 65.50 | 76.27 | 88.58 | 54.27 | 71/89 · 79.8% | 79/95 · 83.2% | USD 0.0335 |
| 17 | JevAny-27B SFT | 68.05 | 95.88 | 86.60 | 86.15 | 48.09 | 87/89 · 97.8% | 93/95 · 97.9% | USD 0.0537 |
| 18 | Visual-Jev 4B Answer-SFT | 67.69 | 79.71 | 81.44 | 86.49 | 49.02 | 76/89 · 85.4% | 87/95 · 91.6% | USD 0.0501 |
| 19 | JevAny-27B RLCR | 67.61 | 95.88 | 87.21 | 86.11 | 47.93 | 87/89 · 97.8% | 93/95 · 97.9% | USD 0.0544 |
| 20 | JPT-0.8Bkirp / llm2jev | 67.49 | 59.00 | 79.00 | 88.56 | 54.43 | 71/89 · 79.8% | 74/95 · 77.9% | USD 0.0330 |
| 21 | Surogate Rune 26B-A4B v3 | 66.69 | 89.70 | 87.31 | 85.99 | 47.92 | 84/89 · 94.4% | 90/95 · 94.7% | USD 0.0544 |
| 22 | OmniJev 0.8Btinnel123 | 66.39 | 57.39 | 73.85 | 88.96 | 55.43 | 74/89 · 83.1% | 71/95 · 74.7% | USD 0.0306 |
| 23 | shisa-de-1 | 66.16 | 86.33 | 89.72 | 85.93 | 47.81 | 83/89 · 93.3% | 88/95 · 92.6% | USD 0.0549 |
| 24 | Reflex 4Breleased stable configuration | 64.10 | 79.61 | 80.73 | 86.06 | 47.96 | 81/89 · 91.0% | 84/95 · 88.4% | USD 0.0543 |
| 25 | JPT-4Bkirp / llm2jev | 62.84 | 89.16 | 93.72 | 85.41 | 46.52 | 85/89 · 95.5% | 89/95 · 93.7% | USD 0.0606 |
| 26 | Gevva E4Btext checkpoint, image input | 62.68 | 60.30 | 59.62 | 87.42 | 52.57 | 71/89 · 79.8% | 75/95 · 78.9% | USD 0.0381 |
| 27 | Jev-Vision 8BSeanLiu | 62.17 | 88.40 | 87.68 | 85.33 | 46.60 | 84/89 · 94.4% | 89/95 · 93.7% | USD 0.0602 |
| 28 | djev-spark NVFP4 | 59.72 | 76.79 | 91.38 | 83.47 | 46.34 | 79/89 · 88.8% | 83/95 · 87.4% | USD 0.0615 |
| 29 | djev-dev BF16 | 55.52 | 77.77 | 87.86 | 82.97 | 45.05 | 82/89 · 92.1% | 82/95 · 86.3% | USD 0.0679 |
| 30 | djev-distill-v4tarsur385 | 55.01 | 77.77 | 86.14 | 82.92 | 44.95 | 82/89 · 92.1% | 82/95 · 86.3% | USD 0.0684 |
| 31 | diffusiongemma-26b djev v10 step160snowicarus | 54.32 | 79.61 | 86.46 | 82.79 | 44.61 | 81/89 · 91.0% | 84/95 · 88.4% | USD 0.0702 |
| 32 | JPT-9Bkirp / llm2jev | 53.73 | 95.34 | 90.60 | 83.99 | 43.52 | 88/89 · 98.9% | 92/95 · 96.8% | USD 0.0763 |
| 33 | Visual-Jev genericzero-shot Qwen3.5-4B baseline | 53.11 | 81.99 | 86.36 | 83.57 | 44.04 | 79/89 · 88.8% | 87/95 · 91.6% | USD 0.0733 |
| 34 | CUA-S1 4B | 52.20 | 83.29 | 80.56 | 84.21 | 43.90 | 79/89 · 88.8% | 88/95 · 92.6% | USD 0.0741 |
| 35 | JEVisiondivyanshx11, visual route | 51.35 | 60.53 | 63.80 | 85.10 | 45.93 | 73/89 · 82.0% | 74/95 · 77.9% | USD 0.0635 |
| 36 | Gevva E2B multimodal | 50.39 | 45.66 | 68.49 | 88.21 | 54.51 | 62/89 · 69.7% | 69/95 · 72.6% | USD 0.0328 |
| 37 | Autoloops – Gemma 4 31B ITAPI | 49.44 | 93.82 | 87.91 | 73.27 | 42.73 | 86/89 · 96.6% | 92/95 · 96.8% | USD 0.0811 |
| 38 | Standard One 8B | 49.10 | 67.03 | 80.48 | 83.65 | 43.69 | 73/89 · 82.0% | 79/95 · 83.2% | USD 0.0754 |
| 39 | imajev 9B | 47.48 | 93.82 | 93.34 | 82.92 | 41.35 | 86/89 · 96.6% | 92/95 · 96.8% | USD 0.0902 |
| 40 | Standard One 3B | 44.79 | 45.88 | 79.13 | 85.27 | 47.30 | 64/89 · 71.9% | 68/95 · 71.6% | USD 0.0571 |
| 41 | Winnow-12BQ8_0 + F16 vision projector | 44.49 | 95.34 | 96.44 | 81.88 | 40.15 | 88/89 · 98.9% | 92/95 · 96.8% | USD 0.0988 |
| 42 | Qevi-2BMeerDevelopment | 36.72 | 41.00 | 52.57 | 88.52 | 53.99 | 61/89 · 68.5% | 66/95 · 69.5% | USD 0.0342 |
| 43 | Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj | 19.89 | 96.64 | 93.52 | 72.38 | 29.29 | 88/89 · 98.9% | 93/95 · 97.9% | USD 0.2275 |
| 44 | GPT-6 Lunalow reasoning effort Setting: OpenRouter reasoning.effort=low Cost receipts cover 95.1% of calls API | 13.64 | 87.00 | 92.79 | 61.95 | 25.68 | 89/89 · 100.0% | 85/95 · 89.5% | USD 0.2712 |
| 45 | GPT-5.6 Luna Cost receipts cover 99.5% of calls API | 11.43 | 98.70 | 96.54 | 61.76 | 23.73 | 89/89 · 100.0% | 94/95 · 98.9% | USD 0.3451 |
| 46 | Gemini 3.1 Flash LiteAPI | 10.89 | 100.00 | 92.93 | 61.79 | 23.31 | 89/89 · 100.0% | 95/95 · 100.0% | USD 0.3601 |
| 47 | PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor) | 0.77 | 9.00 | 35.23 | 89.09 | 55.73 | 48/89 · 53.9% | 49/95 · 51.6% | USD 0.0299 |
| 48 | Gemini 3.8 Flash0.0 · gated by Cost (USD 1.58 per 1,000 decisions)API | 0.09 | 98.70 | 70.41 | 57.60 | 4.01 | 89/89 · 100.0% | 94/95 · 98.9% | USD 1.5839 |
| 49 | OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported) | 0.00 | 75.93 | 0.00 | 79.53 | 39.98 | 83/89 · 93.3% | 80/95 · 84.2% | USD 0.1002 |
| 50 | OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported) | 0.00 | 88.40 | 0.00 | 79.45 | 40.00 | 84/89 · 94.4% | 89/95 · 93.7% | USD 0.1000 |
djev-spark sealed photo result: 83/95 sealed decisions · 87.4%. It saw the public promo images in an earlier inference-only video run, with no training; this sealed score is the independent measurement for it.
The split is 228 public / 456 sealed (33.3% / 66.7%). The public part is unchanged. The sealed part is 123 never-exposed v0.1 items plus 333 fresh items from our private synthetic rotation pool, and all 50 systems were re-run on the fresh items with their original settings. 93 legacy items are retired and not scored. Split counts and hashes were frozen and posted before any system saw a fresh item.
Public / sealed split
228 / 456
33.3% public · 66.7% sealed · was 228 / 216
Licensed real-source items
201 items
139 public · 62 sealed · 93 retired
Fresh sealed items
333 items
Our own synthetic renders and photos · 299 core · 34 everyday photos
Synthetic share
70.6%
483/684 scored items are our own synthetic content
Public means an item was already exposed anywhere. That includes all 79 Mind2Web-derived items because Kev's training data overlaps Mind2Web; the 89 promo photos shown in videos; the 8 example cards on this page and in the status video; and 61 source rows in the public site repository since the 21 Sep preview. The split also counts 1 image asset already present in that repository. Every scored item that has never been exposed is sealed.
Sealed is 123 never-exposed v0.1 items plus 333 fresh items drawn from our private synthetic rotation pool (documents, charts, inventory and safety scenes, and everyday photos). The draw was stratified by family and difficulty with a seeded draw, and the split counts and hashes were frozen before any system saw a fresh item. Computer Use and Browser Use pool items were excluded because those tracks stay separate. Existing item-level outputs are reused for the public and kept sealed items; every system was run on the fresh items with the same code, pinned revisions, prompts and settings as its original run.
Retired. 93 items (ScreenSpot 45, ScreenSpot-Pro 31, Android-in-the-Wild 17) were moved from public to sealed on 24 Sep while their inputs and every system's predictions sat outside the sealed store, so they cannot count as an unseen holdout. They are not scored and not relabelled.
| Family | Public | Sealed | Retired |
|---|---|---|---|
| Android-in-the-Wild (AITW_Single mirror) | 0 | 11 | 17 |
| Everyday photo | 89 | 95 | 0 |
| FinQA | 20 | 0 | 0 |
| Geometry3K | 19 | 0 | 0 |
| Multimodal-Mind2Web | 79 | 0 | 0 |
| Pool: charts | 0 | 79 | 0 |
| Pool: documents | 0 | 102 | 0 |
| Pool: inventory | 0 | 61 | 0 |
| Pool: safety inspection | 0 | 57 | 0 |
| ScreenSpot | 20 | 29 | 45 |
| ScreenSpot-Pro | 1 | 22 | 31 |
| Total | 228 | 456 | 93 |
Two planned tracks extend the image benchmark to computer and browser interactions.
Computer Use
400 items
147 public · 253 sealed
Public decision types
Sealed decision types
Sealed origin: original synthetic local fixtures
Browser Use
400 items
240 public · 160 sealed
Public decision types
Sealed decision types
Sealed origin: original synthetic local fixtures
Cross-track sealing rule. A public-origin Computer Use or Browser Use row that shares a source row or screenshot with a sealed core item is sealed too.
Kev / Mind2Web flag. 133 Mind2Web-derived Browser Use rows are public-only and reported separately for Kev, whose training data overlaps Mind2Web.
Not measured yet — no scores. Neither track has a reviewed system or reusable output. Computer Use has no reviewed candidate in the current queue. Browser Use has one candidate requested, not yet evaluated: kev-0.6b-browser-use. Its public, ungated HF adapter is pinned at 08414b0001f6eba32bc69372abfeac5b74718ac0 (model card license: Apache-2.0), with base jaredpalmer/kev-0.6b at dece6dba8d43f0f7ded45e9f5b9df12474d90843. The model takes textual DOM state and uses a pointer head; its serving path and safe head loader are not independently reviewed. Its Mind2Web-derived examples remain public-only. Both tracks still lack a ready matched-family split and measurements, so they remain separate, requested, and unscored outside the core ranking.
Intelligence. Accuracy counts missing, invalid and unparseable answers as wrong. Each part is chance-corrected against its own average chance rate, then combined as 35% public and 65% sealed. Calibration uses the same weights.
Matched-family overfit penalty. The gap is public accuracy minus sealed accuracy within families that have at least 10 items on both sides: ScreenSpot and Everyday photo. If that matched gap is above 15 percentage points, Intelligence is multiplied by max(0, 1 − (gap − 15)/100). The same rule applies to every system. The raw overall gap is shown in the data but does not affect the score.
Calibration. Ten-bin top-label ECE is scaled by valid probability coverage. Label-only output receives zero calibration. OpenJev's NLI entailment values select an answer but are not treated as categorical probabilities.
Speed and cost. Speed uses whole-call p50 and p95 latency; local latency uses the v1.4 2× plus 0.15-second adjustment. Hosted unit cost uses returned per-call usage receipts; missing receipts are not zero-filled, and the Cost axis is scaled by receipt coverage. Retry costs are tracked separately. Wity-1 uses an estimated base-model price of USD 0.15 per million input tokens and USD 1.00 per million output tokens, applied to its measured 120,542 input and zero output tokens. Florian chose this basis on 29 Sep because Wity's younger public tariff is below the base-model reference. The API does not quantify image or thinking work, so this estimate may understate full image-compute cost. Local cost uses measured GPU seconds at the recorded per-system GPU-hour rate and excludes loading, downloads, build, and idle time. For self-hosted systems, the original run's timings are combined with the fresh-item run on the same GPU type.
Composite and gates. These rules are unchanged. The four axes use an equal-weight harmonic mean, followed by the Jev-class Intelligence, Speed, and Cost gates below 50. Gemini 3.8 Flash's high raw accuracy but near-zero composite reflects its measured cost and the Cost gate; label-only systems have zero Calibration under the inherited convention.
Difficulty balance. The split follows exposure, not a stratified draw, so the parts differ in family mix: browser actions (Mind2Web), chart questions (FinQA) and geometry are public-only, while ScreenSpot-Pro, Android-in-the-Wild and the fresh pool families are sealed-only. This is why the overfit penalty compares only matched families (ScreenSpot and Everyday photo).
Fresh-item difficulty. The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems. Scores on this split are therefore not comparable with the earlier 228/216 preview.
Exposure. GPT-6 Luna and Gemini 3.8 Flash previously saw public promo-photo candidates and sealed-photo candidates in stateless label-check calls, including candidates later dropped. The checks showed no gold; human gold-blind adjudication decided inclusion. GPT-6 Luna is the saved low-reasoning-effort setting. Four OpenRouter systems (GPT-6 Luna low, GPT-5.6 Luna, Gemini 3.1 Flash Lite and Gemini 3.8 Flash) received sealed images and questions through the requested no-retention route with provider fallback disabled. Gemma 4 31B used the Autoloops endpoint. Wity-1 was run on 29 Sep through the hosted Wity SystemOne endpoint with bearer authentication; no no-retention claim is made for Wity or Autoloops. These six hosted rows carry the API flag. Autoloops – Gemma 4 31B IT: results on the original items are reused from the completed Autoloops measurement; the 333 fresh sealed items were run on 25 Sep with the original Autoloops runner. Returned per-request token usage is costed at the published Gemma 4 31B rates. djev-spark saw the 100 public promo photos in the earlier inference-only video run, with no training. Its score on the 95 sealed photo items is the independent measure. Local systems ran without network, credentials, or gold maps. The fresh pool items were authored with Claude and reviewed by OpenAI Codex as a blind critic, and the pool photos were generated with an OpenAI image model. This is disclosed for the two OpenAI rows. Self-hosted systems ran the fresh items on one rented GPU pod in offline containers with read-only inputs.
The ranking covers 50 measured configurations. Wity-1 completed a full 684-decision API run. Its Cost axis uses the Qwen3.6-35B-A3B market reference applied to returned usage, by the 29 Sep price decision; the server build remains under author review. Candidates below have no score unless listed in the ranking. Requested rows remain visible with the exact access or review blocker; exclusions describe the reviewed interface, license or duplicate status.
| Candidate | Source revision | Access | Status | Reason |
|---|---|---|---|---|
| CUA-S1-4B-0.2 multimodal adapter | HF 16818868b | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#23 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Visual Jev 4B Answer-SFT | HF 7a3f1bb0d | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#5 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| NeoHorse-Jev-4B | HF 56c36ae62 | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#4 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Jevify Qwen3-VL-2B Tier 2 | HF 46e8e72c2 | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#25 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Standard One 3B | HF c0d23877e | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#37 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Standard One 8B | HF 5d1285dd7 | Public, ungated; pinned revision resolves | included in v0.1.4 ranking (#24 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Sage1 / Levanto Sage | Hosted image | API requires SAGE_API_KEY; no configured key or free grant | requested, not yet evaluated | No-cost access is unavailable. No paid quota was purchased. |
| djev-distill-v4 | 26B-class ch | No checkpoint staged | included in v0.1.4 ranking (#28 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| yah01/vjev-vision | HF 2fa8b58e4 | Public, ungated | included in v0.1.4 ranking (#31 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| WIlfLin/JEV-Qwen3.8-Flash-Next-Linear-Runtime | HF 1847787ff | Public model metadata; upstream weights about 184 GB | requested, not yet evaluated | License permission and a viable reviewed multi-GPU runtime remain unresolved; no weights were downloaded. |
| Visual-Jev generic Qwen3.5-4B scorer | GitHub 2d68c | Public repository | included in v0.1.4 ranking (#27 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| divyanshx11/JEVision | HF 324698caa | Public, ungated | included in v0.1.4 ranking (#41 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| SeanLiu/Jev-Vision 8B | Exact infere | Checkpoint page is public | included in v0.1.4 ranking (#20 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Fr0zencr4nE/jev-spatial | HF 5727eade6 | Metadata public; previous anonymous weight-file inventory returned HTTP 401 | included in v0.1.4 ranking (#11 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Sarashina 2.2 vision JEV mmproj 3B | HF 09ce275e3 | Public, ungated | requested, not yet evaluated | The vision-language projection has no reviewed general typed-choice adapter/runtime. |
| OhtaMan Gemma 4 E2B IT choice-64 | HF 3d403b990 | Public, ungated | requested, not yet evaluated | Image-text metadata is present, but its actual choice interface and inference code are unreviewed. |
| MetaSK-Jev 4B policy mix | HF ea20fe85b | Public, ungated | requested, not yet evaluated | Image-text metadata is present, but its actual image path and typed-choice runtime are unreviewed. |
| Jevify Gemma 4 E4B | HF a6b5a716f | Public, ungated | included in v0.1.4 ranking (#36 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Jevify Gemma 4 26B-A4B | HF d4c0d1d45 | Public, ungated | included in v0.1.4 ranking (#16 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| JevAny-27B-SFT | HF ad7b48b70 | Public, ungated | included in v0.1.4 ranking (#18 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| JevAny-27B-RLCR | HF 078883b2e | Public, ungated | included in v0.1.4 ranking (#19 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| LFM2.5-VL-3B-Decision-NVFP4 | HF 1ff8bf556 | Public, ungated | requested, not yet evaluated | Image/video input is reported, but license terms, decision-head behavior and local runtime are unresolved. |
| yeyan00/Jev-Decision | GitHub eac9d | Public repository; exact checkpoint access unknown | requested, not yet evaluated | The checkpoint, image interface and inference runtime are not pinned or reviewed. |
| Bonsai-Llama-Jev submission (measured as Bonsai-2-27B v2) | kyr0/Bonsai- | Public pinned submission and model artifacts; completed item-level outputs reused from the Bonsai measurement job | included in v0.1.4 ranking (#42 of 50) | Bonsai-Llama-Jev is retained from the live ImageJevBench roster; see its ranking row above. |
| AutoJev-27B | GitHub denis | Public, ungated; Apache-2.0 weights / MIT code | included in v0.1.4 ranking (#9 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| CLM-8B | Contrastive- | Public checkpoint; text encoder | excluded from image ranking | The reviewed Engine.answer(state, {decision: question}) adapter and frozen Qwen3-8B encoder have no image argument or image processor. |
| Eikos 4B / 27B | Reviewed Let | Public checkpoint; source reviewed | excluded from image ranking | LetterAdapter.dist(state_text, question, options) receives text only; conditional-generation architecture selection is not image input. |
| Main Jev / TypeSafe | Official doc | Official API/interface; no native image field | excluded from image ranking | The documented state/decision interface accepts text, objects and arrays, with no supported image input. A wrapper field is insufficient. |
| OpenJev-Vision | GitHub b83be | Public repository | excluded from core ranking | The published fixed classifier/head does not expose the general typed dynamic-option interface used by the benchmark. |
| hr98w/jev-visual | GitHub 4382b | Public repository | excluded as a duplicate wrapper | A request/runtime wrapper, not a distinct checkpoint or decision head. |
| Jev-Omni reuploads and quantizations | Candidate sw | Public reuploads | excluded as duplicates | No distinct model family/runtime is established beyond the already measured Jev-Omni row. |
| joyfox/Qwen3.5-0.8B-JEV | Candidate sw | Public checkpoint | excluded from image ranking | The decision artifact removes the vision tower; text-only. |
| Hanno-Labs/bosun-v3.1 0.6B / 1.7B | Candidate sw | Public checkpoint | excluded from image ranking | Qwen3 text decision models; no image path documented in the reviewed release. |
| autotrust/JEV | Candidate sw | Public source | excluded from image ranking | Typed text/JSON decisions only; no image input documented in the reviewed interface. |
| Programalyst/realtime-vision-decision-agent | Candidate sw | Public repository | excluded as an application wrapper | Combines vision detection with Jev but is not a distinct model checkpoint. |
| JPT-0.8B | Pinned measu | Public checkpoint; non-commercial terms | included in v0.1.4 ranking (#35 of 50) | Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above. |
| JPT-4B | Pinned measu | Public checkpoint; non-commercial terms | included in v0.1.4 ranking (#6 of 50) | Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above. |
| JPT-9B | Pinned measu | Public checkpoint; non-commercial terms | included in v0.1.4 ranking (#12 of 50) | Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above. |
| Qevi-2B | Pinned measu | Public checkpoint; non-commercial terms | included in v0.1.4 ranking (#43 of 50) | Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above. |
| OmniJev 0.8B | Pinned measu | Public aggregate row | included in v0.1.4 ranking (#39 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| OmniJev 2B | Pinned measu | Public aggregate row | included in v0.1.4 ranking (#22 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| OmniJev 4B | Pinned measu | Public aggregate row | included in v0.1.4 ranking (#13 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Imajev-4B | Imajev-4B fa | Public request model | included in v0.1.4 ranking (#1 of 50) | Remeasured on the unchanged v0.1 method and split with the pinned server and adapter, one rotation, --fast, and --merge-lora. Only aggregate results are published. |
| Glance frozen Qwen3-VL-4B | yoheinakajim | Frozen checkpoint; requested by Yohei | included in v0.1.4 ranking (#8 of 50) | Retained from the reviewed ImageJev aggregate. The hold applies only to the separate Qwen3-VL-2B CUDA speedlab path. |
| Laya Vision | thaitea/laya | Customer-requested evaluation | held at author's request | Not ranked or published while the author explores other options. |
| Glance speedlab Qwen3-VL-2B | glance.yohei | Apple-only MLX path; measured CUDA port differs | held pending author discussion | Not ranked or published until Yohei is asked about the separate CUDA port. |
| Winnow-12B | EldanRing/wi | Reviewed source pin; independently scored ImageJev handoff | included in v0.1.4 ranking (#34 of 50) | Independent recomputation matched 21/21 metrics with zero delta. The pinned source and author build.py/runtime lock were used, but this run used an isolated server image rather than the separate multi-model load-test Docker context. Medium dependency reproducibility: apt packages resolved at build time and pip wheels lack hashes. All per-file hashes and image membership passed; a separate image-tree digest was not compared with the aggregate receipt (low assurance limitation). |
| Jev-Omni | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#3 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| imajev 2B | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#7 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Mapika decider-2b-vision BF16 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#10 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Reflex 4B (released stable configuration) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#14 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| OmniJev-Qwen3.5-9B-v4 (tzcfly) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#15 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Surogate Rune 26B-A4B v3 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#17 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| shisa-de-1 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#21 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| imajev 9B | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#26 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| djev-spark NVFP4 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#29 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| diffusiongemma-26b djev v10 step160 (snowicarus) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#30 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Autoloops – Gemma 4 31B IT | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#32 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| djev-dev BF16 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#33 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Gevva E4B (text checkpoint, image input) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#38 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Gevva E2B multimodal | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#40 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| GPT-6 Luna (low reasoning effort) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#44 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| GPT-5.6 Luna | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#45 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Gemini 3.1 Flash Lite | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#46 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| PlayJev 0.8B | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#47 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Gemini 3.8 Flash | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#48 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| OpenJev 4B NLI v2 (official image-premise path) | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#49 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| OpenJev 4B NLI v5 | Frozen v0.1. | Public aggregate row | included in v0.1.4 ranking (#50 of 50) | Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above. |
| Wity-1 | Wity product | Hosted Wity SystemOne API; Cost estimated from Qwen3.6-35B-A3B market reference | included in v0.1.4 ranking (#2 of 50) | Full 684-decision run with valid probabilities and complete input/output usage; base-model reference applied to measured tokens. Server build under author review. |
Image JevBench is a separate benchmark from the text-only JevBench Score. Sealed item-level content remains private. Public/sealed item counts, accuracy, track and score breakdowns are aggregates. Results describe these exact tested configurations and do not establish absence from model training data.