Favur Evals
favur.dev

Model · xAI

xAI Grok 4.5

4 appearances · meta-score 72.5/100 · last seen 2026-07-22 · scored under rubric v15

Pooled composite across 4 primary runs: 62.8 ± 2.5 (95% CI)

Agent role scores

Run history (primary model)

Runs where xAI Grok 4.5 was the primary model, oldest first
RunSoWDateComposite
Grok 4.5circlesJul 202661.0
Grok 4.5circlesJul 202663.2
Grok 4.5circlesJul 202662.3
Grok 4.5circlesJul 202664.7

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.