Favur Evals
favur.dev

Model · Anthropic

Anthropic Claude Sonnet 4.6

2 appearances · meta-score 43.2/100 · last seen 2026-07-03 · scored under rubric v15

Pooled composite across 2 primary runs: 43.4 ± 13.3 (95% CI) provisional: n < 3, treat as unsettled

Agent role scores

Per-agent-role scores for Anthropic Claude Sonnet 4.6
Agent roleScore / 10
Scout Agent6.47
Sprint-Plan Agent6.17
Code Agent4.95
Test Agent4.78
Pseudocode Agent4.63
Build Agent4.06
Code-Review Agent3.75
Develop Agent3.57
Orchestrator3.20
Sprint-Review Agent3.15

Run history (primary model)

Runs where Anthropic Claude Sonnet 4.6 was the primary model, oldest first
RunSoWDateComposite
Claude Sonnet 4.6circlesJun 202642.3
Claude Sonnet 4.6circlesJul 202644.4

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.