Agent role · weight 3/100 · rubric v15
📐 Spec
Decision-artifact producer — concrete values only (no TBD/hedging), with a self-check step.
- 6.68Pooled mean / 10
- ± 0.2795% CI
- 33Appearances (n)
- 14Models tested
What it's scored on
- Output quality
- Quality of the decision artifact it wrote and deliverables.
- Efficiency
- Token volume, output per request, cost, context efficiency, request-error and hallucination rates.
Full scoring definition in the methodology → · Sample transcript
Model leaderboard
| # | Model | Score / 10 | Coverage | n | Last seen |
|---|---|---|---|---|---|
| 1 | Meta Muse Spark 1.1 | 7.81 | 96% | 3 | 2026-07-19 |
| 2 | Xiaomi Mimo v2.5 | 7.31 | 96% | 2 | 2026-07-26 |
| 3 | Google Gemini Flash 3.5 † | 7.11 | 76% | 1 | 2026-07-08 |
| 4 | DeepSeek Deepseek Flash v4 | 7.01 | 79% | 4 | 2026-07-22 |
| 5 | Moonshot AI Kimi Code k2.7 † | 6.92 | 76% | 1 | 2026-07-08 |
| 6 | xAI Grok 4.5 | 6.90 | 91% | 4 | 2026-07-22 |
| 7 | Google Gemini Flash 3 (preview) | 6.86 | 72% | 4 | 2026-07-09 |
| 8 | OpenAI Gpt Terra 5.6 † | 6.72 | 76% | 1 | 2026-07-12 |
| 9 | DeepSeek Deepseek Pro v4 † | 6.60 | 76% | 1 | 2026-07-05 |
| 10 | Google Gemini Flash 3.6 | 6.31 | 96% | 4 | 2026-07-24 |
| 11 | Anthropic Claude Sonnet 5 | 6.16 | 86% | 4 | 2026-07-23 |
| 12 | Google Gemini Flash Lite 3.5 † | 6.09 | 96% | 1 | 2026-07-21 |
| 13 | Google Gemini Pro 3.1 (preview) † | 5.81 | 68% | 1 | 2026-06-28 |
| 14 | Google Gemini Flash Lite 3.1 † | 5.12 | 63% | 2 | 2026-07-15 |
† low confidence — sparse role or low data coverage (renormalized, never zero-filled).
Per-run history (newest first)
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.