Bespoke Nimble 9B (Bespoke Labs)
Jev rebuild · by Bespoke Labs
JevBench v1.4.2.2 score
18.7
Rank #62 of 91 ranked systems.
44.6 points behind Jev 1.13.0's 63.3.
Published axes
- intelligence
- 46.3
- calibration
- 56.4
- speed
- 78.7
- cost
- 33.4
Against Jev 1.13.0
Four radars compare this fixed pair across the score axes, accuracy per tier, and accuracy by family on the hard tier and sealed set. Further out is better on every spoke.
- A: Bespoke Nimble 9B — Jev rebuild · Score 18.7 (#62)
- B: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#4)
The four score axes
Accuracy per tier, incl. sealed
Current question set by family (hard + sealed)
Sealed set by family
Availability and evidence
- Openness
- Code and weights marked open in the published row
- License note
- Apache-2.0 (weights); repository without a licence file as of 19 Sep
- Cost evidence
- estimated; the board’s row disclosure contains the published basis.
- Endpoint condition
- our RunPod GPU (A40 48 GB, Canada), reached over the internet from Germany
- Note on this row
- Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- Published source
- https://github.com/bespokelabsai/nimble
From the public v1.4.2.2 aggregate. Scores and ranks can change when a new release is published.
Read the full board and published method. The overall score is a composite, not raw accuracy.