Agent roles
The composite is the weight-averaged mean of these role scores (weights sum to 100). Pooled mean ± 95% CI across every appearance; small-n roles are provisional.
| Role | Weight | Pooled mean / 10 | 95% CI | n |
|---|---|---|---|---|
| Code Review | 14 | 6.26 | ± 0.31 | 75 |
| Code | 14 | 5.74 | ± 0.18 | 75 |
| Develop | 14 | 6.50 | ± 0.34 | 75 |
| Sprint Plan | 9 | 6.77 | ± 0.19 | 75 |
| Pseudocode | 9 | 5.54 | ± 0.19 | 75 |
| Sprint Review | 9 | 6.64 | ± 0.43 | 72 |
| Director | 6 | 6.74 | ± 1.77 | 6 |
| Architect | 6 | 7.11 | ± 0.44 | 35 |
| Orchestrator | 4 | 6.39 | ± 0.35 | 75 |
| Test | 4 | 6.11 | ± 0.29 | 73 |
| Build | 3 | 5.86 | ± 0.31 | 72 |
| Platform | 3 | 5.60 | ± 0.27 | 44 |
| Spec | 3 | 6.68 | ± 0.27 | 33 |
| Scout | 2 | 6.62 | ± 0.18 | 71 |
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.