Favur Evals
favur.dev

Model · Moonshot AI

Moonshot AI Kimi k2.6

3 appearances · meta-score 70.0/100 · last seen 2026-07-17 · scored under rubric v15

Pooled composite across 3 primary runs: 55.5 ± 17.4 (95% CI)

Agent role scores

Per-agent-role scores for Moonshot AI Kimi k2.6
Agent roleScore / 10
Sprint-Review Agent8.17
Code-Review Agent7.79
Develop Agent7.73
Test Agent7.32
Orchestrator7.28
Build Agent7.13
Scout Agent7.04
Sprint-Plan Agent7.03
Code Agent5.96
Pseudocode Agent5.58
Platform Agent4.44

Run history (primary model)

Runs where Moonshot AI Kimi k2.6 was the primary model, oldest first
RunSoWDateComposite
Kimi k2.6circlesJun 202653.1
Kimi k2.6circlesJul 202649.9
Kimi k2.6circlesJul 202663.3

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.