Model · xAI
xAI Grok 4.5
4 appearances · meta-score 72.5/100 · last seen 2026-07-22 · scored under rubric v15
Pooled composite across 4 primary runs: 62.8 ± 2.5 (95% CI)
- $103.11Total cost
- 151,511,297Tokens
- 85.4%Cache hit
- 2,858Requests
- 0.14%Failure rate
- 1.18Tools/response
- 1,469,402Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Sprint-Review Agent | 8.71 |
| Orchestrator | 8.14 |
| Develop Agent | 8.11 |
| Build Agent | 7.80 |
| Sprint-Plan Agent | 7.58 |
| Scout Agent | 7.36 |
| Spec Agent | 7.19 |
| Code-Review Agent | 7.02 |
| Architect | 6.99 |
| Code Agent | 6.75 |
| Platform Agent | 6.31 |
| Test Agent | 6.22 |
| Pseudocode Agent | 5.56 |
Run history (primary model)
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.