Model · Anthropic
Anthropic Claude Sonnet 4.6
2 appearances · meta-score 43.2/100 · last seen 2026-07-03 · scored under rubric v15
Pooled composite across 2 primary runs: 43.4 ± 13.3 (95% CI) provisional: n < 3, treat as unsettled
- $208.10Total cost
- 99,320,992Tokens
- 42.7%Cache hit
- 1,518Requests
- 0.20%Failure rate
- 0.99Tools/response
- 477,276Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Scout Agent | 6.47 |
| Sprint-Plan Agent | 6.17 |
| Code Agent | 4.95 |
| Test Agent | 4.78 |
| Pseudocode Agent | 4.63 |
| Build Agent | 4.06 |
| Code-Review Agent | 3.75 |
| Develop Agent | 3.57 |
| Orchestrator | 3.20 |
| Sprint-Review Agent | 3.15 |
Run history (primary model)
| Run | SoW | Date | Composite |
|---|---|---|---|
| Claude Sonnet 4.6 | circles | Jun 2026 | 42.3 |
| Claude Sonnet 4.6 | circles | Jul 2026 | 44.4 |
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.