jev-tmmluplus-eval
Evaluate TypeSafe AI's Jev (System One Model) on TMMLU+ v1.1 — four-way choice via the API's own response schema, 100% parse rate by construction
Unverified: Jev confidence 0.61; publishing needs more than 0.65. all unverified
- Author
- lianghsun
- Category
- benchmark
- Primitives
- choice
- Source
- github
Classified by jev-1.13.0. Description taken from the source, not generated.