Model comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.
Jev was compared with DeepSeek V4 Flash, GPT-5.6 Luna, GPT-4.1 Mini, and Gemma 4B on 616 answer-correctness pairs from JudgeBench and 1,086 groundedness claims from LLM-AggreFact. Each judgment was scored against an objectively verified or human-authored reference label.
1,086 claimshigher is better
616 independent pairshigher is better
Kimi K3 ran through Moonshot and Fireworks in 30 direct latency requests per provider and 60 planned Figma-to-HTML runs per provider. Paired intervals compare matched designs.
55% faster
Loop and Braintrust MCP + Codex each ran five workflows on identical 200-trace projects. Each tool was scored against a handwritten rubric measuring how well it surfaced insights about the traces and cited supporting evidence so traceability could be verified.
22 condition-neutral checks
39% faster
Each skill packages best practices from leading AI teams and researchers into a plain SKILL.md file your coding agent can read. Use them to get started running evals on your data.
Ask it to do what the skill describes, like validating a scorer against human labels or finding hidden failure modes.