Jev Show

jev-tmmluplus-eval

Evaluate TypeSafe AI's Jev (System One Model) on TMMLU+ v1.1 — four-way choice via the API's own response schema, 100% parse rate by construction

Unverified: Jev confidence 0.61; publishing needs more than 0.65. all unverified

Author
lianghsun
Category
benchmark
Primitives
choice
Source
github

Classified by jev-1.13.0. Description taken from the source, not generated.