Model · Moonshot AI
Moonshot AI Kimi k2.6
3 appearances · meta-score 70.0/100 · last seen 2026-07-17 · scored under rubric v15
Pooled composite across 3 primary runs: 55.5 ± 17.4 (95% CI)
- $77.11Total cost
- 128,253,181Tokens
- 48.9%Cache hit
- 2,549Requests
- 0.16%Failure rate
- 1.17Tools/response
- 1,663,293Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Sprint-Review Agent | 8.17 |
| Code-Review Agent | 7.79 |
| Develop Agent | 7.73 |
| Test Agent | 7.32 |
| Orchestrator | 7.28 |
| Build Agent | 7.13 |
| Scout Agent | 7.04 |
| Sprint-Plan Agent | 7.03 |
| Code Agent | 5.96 |
| Pseudocode Agent | 5.58 |
| Platform Agent | 4.44 |
Run history (primary model)
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.