Models
21 models, listed alphabetically. The meta-score is the role-weighted composite built from each model's best per-role performances. The pooled composite is the mean across the model's primary runs with a 95% confidence interval (Student-t) — entries with fewer than3 primary runs are provisional † and should be read as unsettled.
† provisional — fewer than 3 primary runs; the interval is wide or undefined and the ranking is not settled. Where two models' intervals overlap, their order is not statistically separated.
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.