Trends
With repeated runs under a versioned rubric, drift is the most valuable signal in the dataset. Everything below is derived at build time from the published composites — the same numbers on every run page — and every chart has its table right beside it.
Cohort drift
Is the whole cohort improving, or is the rubric getting easier? Both readings, published: the least-squares slope of composite against run sequence is+0.222 points per runacross 75 runs.Every run was scored under rubric v15, so measurement change cannot explain any of it — whatever drift exists is the models and the rosters.
| Month | Runs | Mean composite |
|---|---|---|
| 2026-06 | 11 | 50.9 |
| 2026-07 | 64 | 61.2 |
Automated regression detection
A run is flagged when its composite falls outside the 95% confidence interval pooled from that model's prior runs (at least two priors required). Moves in both directions are published — improvements are flagged with the same rule as regressions.
| Model | Run | Date | Composite | Prior 95% interval | Direction |
|---|---|---|---|---|---|
| DeepSeek Deepseek Flash v4 | circles/deepseek-deepseek-v4-flash-7.22.2026 | 2026-07-22 | 70.1 | 50.4–62.7 (n=5) | improvement |
| Google Gemini Flash 3.6 | the2048/gemini-multi-model-7-22-2026 | 2026-07-22 | 52.6 | 57.0–82.5 (n=3) | regression |
| Google Gemini Flash 3.6 | solarSystem/google-multi-model-7.21.2026 | 2026-07-21 | 63.8 | 65.7–79.6 (n=2) | regression |
| Google Gemini Flash Lite 3.1 | circles/google-gemini-3.1-flash-lite-7.19.2026 | 2026-07-19 | 51.0 | 59.6–68.5 (n=3) | regression |
| Anthropic Claude Sonnet 5 | circles/anthropic-claude-sonnet-5-7.16.2026 | 2026-07-17 | 70.8 | 39.1–55.3 (n=2) | improvement |
| Google Gemini Flash 3.5 | circles/google-multi-model-7.15.2026 | 2026-07-15 | 81.8 | 39.8–65.3 (n=3) | improvement |
| Google Gemini Flash 3.5 | solarSystem/google-multi-model-7.15.2026 | 2026-07-15 | 42.6 | 44.1–83.5 (n=5) | regression |
| DeepSeek Deepseek Flash v4 | circles/deepseek-deepseek-v4-flash-7.10.2026 | 2026-07-10 | 64.0 | 49.8–59.5 (n=4) | improvement |
Composite over time, per model
Models with at least two primary runs, chronological. Full history on each model page.
† provisional — fewer than 3 primary runs. Models with a single run have no trend yet and are omitted here; they appear on the model index.
Version deltas
When a vendor ships a new version, the delta against its predecessor — meta-score, pooled composite, and every agent role both versions handled.
Anthropic Claude Sonnet 4.6 → Anthropic Claude Sonnet 5
Meta-score +23.7 · Pooled composite +12.9 (n=2 → n=5)
| Agent role | Delta |
|---|---|
| Sprint-Review Agent | +5.0 |
| Orchestrator | +3.9 |
| Develop Agent | +3.9 |
| Code-Review Agent | +3.0 |
| Build Agent | +2.6 |
| Test Agent | +1.2 |
| Code Agent | +0.9 |
| Pseudocode Agent | +0.8 |
| Sprint-Plan Agent | +0.2 |
| Scout Agent | +0.0 |
Google Gemini Flash 3 (preview) → Google Gemini Flash 3.5
Meta-score −1.9 · Pooled composite +14.5 (n=6 → n=8)
| Agent role | Delta |
|---|---|
| Develop Agent | +2.4 |
| Orchestrator | +2.1 |
| Spec Agent | −0.0 |
| Platform Agent | −0.3 |
| Code-Review Agent | −0.3 |
| Code Agent | −0.6 |
| Test Agent | −0.7 |
| Build Agent | −0.7 |
| Sprint-Plan Agent | −0.8 |
| Sprint-Review Agent | −0.9 |
| Pseudocode Agent | −1.3 |
| Scout Agent | −1.5 |
| Architect | −2.2 |
Google Gemini Flash 3.5 → Google Gemini Flash 3.6
Meta-score +12.4 · Pooled composite +1.8 (n=8 → n=5)
| Agent role | Delta |
|---|---|
| Architect | +3.4 |
| Code-Review Agent | +3.3 |
| Sprint-Review Agent | +3.1 |
| Build Agent | +1.8 |
| Platform Agent | +1.3 |
| Code Agent | +1.2 |
| Test Agent | +0.6 |
| Scout Agent | +0.3 |
| Develop Agent | +0.0 |
| Pseudocode Agent | −0.0 |
| Orchestrator | −0.1 |
| Spec Agent | −0.5 |
| Sprint-Plan Agent | −0.7 |
Google Gemini Flash Lite 3.1 → Google Gemini Flash Lite 3.5
Meta-score −3.4 · Pooled composite +17.5 (n=4 → n=1)
| Agent role | Delta |
|---|---|
| Build Agent | +1.8 |
| Develop Agent | +0.4 |
| Scout Agent | +0.3 |
| Code-Review Agent | +0.3 |
| Platform Agent | +0.2 |
| Code Agent | +0.1 |
| Orchestrator | −0.2 |
| Test Agent | −0.5 |
| Spec Agent | −0.5 |
| Sprint-Review Agent | −0.8 |
| Pseudocode Agent | −1.6 |
| Sprint-Plan Agent | −2.4 |
Consistency
Which models are consistent and which are erratic across repeat runs — the sample standard deviation of each model's composites, most consistent first. Consistency is a purchasing criterion nobody publishes.
| Model | n | Mean composite | Std deviation |
|---|---|---|---|
| Google Gemini Pro 3.1 (preview) | 2 | 52.5 | 1.20 † |
| Moonshot AI Kimi Code k2.7 | 2 | 60.2 | 1.29 † |
| Stepfun Step Flash 3.7 | 2 | 57.9 | 1.41 † |
| Anthropic Claude Sonnet 4.6 | 2 | 43.4 | 1.48 † |
| xAI Grok 4.5 | 4 | 62.8 | 1.54 |
| DeepSeek Deepseek Pro v4 | 3 | 53.9 | 2.02 |
| Xiaomi Mimo Pro v2.5 | 5 | 62.5 | 2.81 |
| Google Gemini Flash Lite 3.1 | 4 | 60.8 | 6.68 |
| Moonshot AI Kimi k2.6 | 3 | 55.5 | 7.00 |
| DeepSeek Deepseek Flash v4 | 6 | 58.8 | 7.09 |
| Z.ai Glm 5.2 | 3 | 54.3 | 7.38 |
| Xiaomi Mimo v2.5 | 3 | 74.0 | 8.32 |
| Google Gemini Flash 3.6 | 5 | 64.6 | 8.47 |
| Anthropic Claude Sonnet 5 | 5 | 56.3 | 9.80 |
| Google Gemini Flash 3 (preview) | 6 | 48.2 | 11.54 |
| Qwen Plus qwen3.7 | 4 | 68.9 | 11.62 |
| Meta Muse Spark 1.1 | 3 | 68.3 | 12.05 |
| Google Gemini Flash 3.5 | 8 | 62.7 | 14.97 |
Composite over time, per Statement of Work
circles — 59 runs
solarSystem — 13 runs
Statement of Work: solarSystem
| Date | Run | Composite | Rubric |
|---|---|---|---|
| 2026-06-27 | solarSystem/deepseek-deepseek-v4-flash-6.27.2026 | 50.2 | v15 |
| 2026-06-28 | solarSystem/google-gemini-3-flash-preview-6.27.2026 | 28.9 | v15 |
| 2026-06-29 | solarSystem/xiaomi-mimo-v2.5-pro-6.29.2026 | 57.6 | v15 |
| 2026-07-13 | solarSystem/google-gemini-3.1-flash-lite-7.12.2026 | 62.2 | v15 |
| 2026-07-14 | solarSystem/google-multi-model-7.14.2026 | 57.6 | v15 |
| 2026-07-15 | solarSystem/google-multi-model-7.15.2026 | 42.6 | v15 |
| 2026-07-17 | solarSystem/meta-muse-spark-1.1-7.17.2026 | 67.2 | v15 |
| 2026-07-18 | solarSystem/xiaomi-mimo-v2.5-7.18.2026 | 75.9 | v15 |
| 2026-07-19 | solarSystem/stepfun-step-3.7-flash-7.19.2026 | 56.9 | v15 |
| 2026-07-19 | solarSystem/qwen-qwen3.7-plus-7.18.2026 | 54.7 | v15 |
| 2026-07-21 | solarSystem/google-multi-model-7.21.2026 | 63.8 | v15 |
| 2026-07-23 | solarSystem/multi-anthropic-sonnet-gemini-3.5-flash-lite-7.22.2026 | 58.3 | v15 |
| 2026-07-26 | solarSystem/multi-xiaomi-mimo-v2.5-7.26.2026 | 63.2 | v15 |
the2048 — 3 runs
| Date | Run | Composite | Rubric |
|---|---|---|---|
| 2026-06-26 | the2048/deepseek-deepseek-v4-flash-6.26.2026 | 55.6 | v15 |
| 2026-07-19 | the2048/meta-muse-spark-1.1-7.19.2026 | 56.8 | v15 |
| 2026-07-22 | the2048/gemini-multi-model-7-22-2026 | 52.6 | v15 |
Rubric-version impact
Every published score was produced under rubric v15, so there is no measurement change to separate from model change yet. When a rubric revision lands, runs re-scored under both versions will appear here side by side, and the dataset changelog records what changed.
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.