Favur Evals
favur.dev

Model · Moonshot AI

Moonshot AI Kimi Code k2.7

2 appearances · meta-score 67.4/100 · last seen 2026-07-10 · scored under rubric v15

Pooled composite across 2 primary runs: 60.2 ± 11.6 (95% CI) provisional: n < 3, treat as unsettled

Agent role scores

Per-agent-role scores for Moonshot AI Kimi Code k2.7
Agent roleScore / 10
Sprint-Plan Agent7.55
Pseudocode Agent7.48
Sprint-Review Agent7.32
Scout Agent7.22
Spec Agent6.92
Develop Agent6.72
Code-Review Agent6.62
Code Agent6.53
Orchestrator5.90
Test Agent5.84
Platform Agent5.32
Build Agent5.08

Run history (primary model)

Runs where Moonshot AI Kimi Code k2.7 was the primary model, oldest first
RunSoWDateComposite
Kimi Code k2.7circlesJul 202661.1
Kimi Code k2.7circlesJul 202659.3

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.