Favur Evals
favur.dev

Model · Google

Google Gemini Pro 3.1 (preview)

2 appearances · meta-score 58.0/100 · last seen 2026-07-03 · scored under rubric v15

Pooled composite across 2 primary runs: 52.5 ± 10.8 (95% CI) provisional: n < 3, treat as unsettled

Agent role scores

Per-agent-role scores for Google Gemini Pro 3.1 (preview)
Agent roleScore / 10
Test Agent7.20
Sprint-Plan Agent6.80
Pseudocode Agent5.96
Spec Agent5.81
Code-Review Agent5.76
Code Agent5.75
Sprint-Review Agent5.68
Develop Agent5.68
Orchestrator5.24
Architect5.21
Build Agent4.76
Platform Agent4.59

Run history (primary model)

Runs where Google Gemini Pro 3.1 (preview) was the primary model, oldest first
RunSoWDateComposite
Gemini Pro 3.1 (preview)circlesJun 202651.7
Gemini Pro 3.1 (preview)circlesJul 202653.4

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.