Model · Google
Google Gemini Flash 3.6
5 appearances · meta-score 72.3/100 · last seen 2026-07-24 · scored under rubric v15
Pooled composite across 5 primary runs: 64.6 ± 10.5 (95% CI)
- $212.41Total cost
- 417,382,314Tokens
- 68.0%Cache hit
- 9,328Requests
- 0.32%Failure rate
- 0.96Tools/response
- 1,964,948Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Code-Review Agent | 8.66 |
| Architect | 8.36 |
| Sprint-Review Agent | 8.34 |
| Develop Agent | 8.13 |
| Orchestrator | 7.81 |
| Director | 7.41 |
| Platform Agent | 7.01 |
| Build Agent | 6.82 |
| Test Agent | 6.80 |
| Spec Agent | 6.61 |
| Scout Agent | 6.48 |
| Code Agent | 6.22 |
| Sprint-Plan Agent | 6.05 |
| Pseudocode Agent | 4.90 |
Run history (primary model)
| Run | SoW | Date | Composite |
|---|---|---|---|
| Gemini Flash 3.6 | circles | Jul 2026 | 73.2 |
| Gemini Flash 3.6 + Gemini Flash Lite | circles | Jul 2026 | 72.1 |
| Gemini Flash 3.6 + Gemini Flash Lite | solarSystem | Jul 2026 | 63.8 |
| Gemini Flash 3.6 + Gemini Flash Lite | the2048 | Jul 2026 | 52.6 |
| Gemini Flash 3.6 | circles | Jul 2026 | 61.1 |
- ▼ 2026-07-22 scored 52.6, outside the prior 95% interval 57.0–82.5 (n=3) — a statistically notable regression. All flags.
- ▼ 2026-07-21 scored 63.8, outside the prior 95% interval 65.7–79.6 (n=2) — a statistically notable regression. All flags.
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.