Model · Google
Google Gemini Flash Lite 3.1
4 appearances · meta-score 74.3/100 · last seen 2026-07-19 · scored under rubric v15
Pooled composite across 4 primary runs: 60.8 ± 10.6 (95% CI)
- $18.68Total cost
- 177,639,420Tokens
- 75.2%Cache hit
- 4,305Requests
- 0.00%Failure rate
- 0.93Tools/response
- 9,507,461Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Sprint-Review Agent | 9.08 |
| Orchestrator | 8.07 |
| Sprint-Plan Agent | 7.98 |
| Develop Agent | 7.96 |
| Code-Review Agent | 7.40 |
| Architect | 7.28 |
| Pseudocode Agent | 7.23 |
| Test Agent | 7.17 |
| Scout Agent | 6.79 |
| Spec Agent | 6.64 |
| Code Agent | 6.43 |
| Platform Agent | 6.35 |
| Build Agent | 6.00 |
Run history (primary model)
| Run | SoW | Date | Composite |
|---|---|---|---|
| Gemini Flash Lite 3.1 | circles | Jul 2026 | 65.7 |
| Gemini Flash Lite 3.1 | solarSystem | Jul 2026 | 62.2 |
| Gemini Flash Lite 3.1 | circles | Jul 2026 | 64.2 |
| Gemini Flash Lite 3.1 | circles | Jul 2026 | 51.0 |
- ▼ 2026-07-19 scored 51.0, outside the prior 95% interval 59.6–68.5 (n=3) — a statistically notable regression. All flags.
Also in the roster of
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.