Favur Evals
favur.dev

Models

21 models, listed alphabetically. The meta-score is the role-weighted composite built from each model's best per-role performances. The pooled composite is the mean across the model's primary runs with a 95% confidence interval (Student-t) — entries with fewer than3 primary runs are provisional † and should be read as unsettled.

All evaluated models, alphabetical
ModelVendorMeta-scorePooled composite (95% CI)nAppearancesBest atLast seen
Anthropic Claude Sonnet 4.6Anthropic43.243.4 ± 13.3 †22Scout Agent2026-07-03
Anthropic Claude Sonnet 5Anthropic66.956.3 ± 12.255Sprint-Review Agent2026-07-23
DeepSeek Deepseek Flash v4DeepSeek74.058.8 ± 7.466Director2026-07-22
DeepSeek Deepseek Pro v4DeepSeek58.653.9 ± 5.033Sprint-Plan Agent2026-07-05
Google Gemini Flash 3 (preview)Google61.848.2 ± 12.166Scout Agent2026-07-10
Google Gemini Flash 3.5Google59.962.7 ± 12.588Develop Agent2026-07-16
Google Gemini Flash 3.6Google72.364.6 ± 10.555Code-Review Agent2026-07-24
Google Gemini Flash Lite 3.1Google74.360.8 ± 10.644Sprint-Review Agent2026-07-19
Google Gemini Flash Lite 3.5Google71.078.2 (no CI, n < 2) †11Develop Agent2026-07-21
Google Gemini Pro 3.1 (preview)Google58.052.5 ± 10.8 †22Test Agent2026-07-03
Meta Muse Spark 1.1Meta77.568.3 ± 29.933Sprint-Review Agent2026-07-19
Moonshot AI Kimi Code k2.7Moonshot AI67.460.2 ± 11.6 †22Sprint-Plan Agent2026-07-10
Moonshot AI Kimi k2.6Moonshot AI70.055.5 ± 17.433Sprint-Review Agent2026-07-17
OpenAI Gpt Luna 5.6OpenAI64.761.1 (no CI, n < 2) †11Sprint-Plan Agent2026-07-14
OpenAI Gpt Terra 5.6OpenAI65.461.5 (no CI, n < 2) †11Sprint-Plan Agent2026-07-12
Qwen Plus qwen3.7Qwen78.568.9 ± 18.544Code-Review Agent2026-07-22
Stepfun Step Flash 3.7Stepfun66.557.9 ± 12.7 †22Architect2026-07-19
xAI Grok 4.5xAI72.562.8 ± 2.544Sprint-Review Agent2026-07-22
Xiaomi Mimo Pro v2.5Xiaomi70.462.5 ± 3.555Test Agent2026-07-26
Xiaomi Mimo v2.5Xiaomi77.474.0 ± 20.733Sprint-Review Agent2026-07-22
Z.ai Glm 5.2Z.ai69.954.3 ± 18.333Sprint-Review Agent2026-07-17

† provisional — fewer than 3 primary runs; the interval is wide or undefined and the ranking is not settled. Where two models' intervals overlap, their order is not statistically separated.

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.