open-system-one

An independent benchmark of TypeSafe's Jev against open, CPU-only stacks. 10,000 decisions · 4 datasets · Apple M5 Pro · CPU · batch = 1.
One exception, disclosed: Laya's row is Apple MPS, because laya.load() auto-selects it. Its like-for-like CPU latency is 59.5 / 78.0 / 61.9 / 309.0 ms — about 3× higher.

79.3 %typesafe/jev — best macro
78.7 %open cross-encoder, 149M, CPU
0.1 msstatic embedding + 8-float head
$0.284Jev, 10k decisions (measured)

Accuracy vs latency

Mean p50 latency across the four tasks (log scale). Jev's latency is dominated by network round-trip — see the repo's limitations.

Zero-shot — neither side gets labels

Same option text for every stack. Neither condition is neutral: the rewrite moves Jev +0.5 pp, the bi-encoder +6.3 pp, and the cross-encoder −4.3 pp.

With ~2,000 labels per task not comparable to Jev

Which stack is fastest for your option count?

Cross-encoder cost grows linearly with option count; Jev's is flat (381 ms at 2 options, 392 ms at 77). Extrapolated from measured points.

Eight things that did not work