jev-guardrails
Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage.
Unverified: Jev confidence 0.62; publishing needs more than 0.65. Category confidence 0.47; publishing needs more than 0.60. all unverified
- Author
- deepansh-saxena
- Category
- benchmark
- Primitives
- unknown
- Source
- github
Classified by jev-1.13.0. Description taken from the source, not generated.