Favur Evals
favur.dev

Agent roles

The composite is the weight-averaged mean of these role scores (weights sum to 100). Pooled mean ± 95% CI across every appearance; small-n roles are provisional.

Weighted agent roles with pooled scores and confidence intervals
RoleWeightPooled mean / 1095% CIn
Code Review146.26± 0.3175
Code145.74± 0.1875
Develop146.50± 0.3475
Sprint Plan96.77± 0.1975
Pseudocode95.54± 0.1975
Sprint Review96.64± 0.4372
Director66.74± 1.776
Architect67.11± 0.4435
Orchestrator46.39± 0.3575
Test46.11± 0.2973
Build35.86± 0.3172
Platform35.60± 0.2744
Spec36.68± 0.2733
Scout26.62± 0.1871

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.