Agent role · weight 6/100 · rubric v15
🏛️ Architect
Produces the architecture document and ADRs — interfaces, technology decisions and constraints for the phase.
- 7.11Pooled mean / 10
- ± 0.4495% CI
- 35Appearances (n)
- 15Models tested
What it's scored on
- Output quality
- Architecture completeness (components documented, ADRs recorded, document written) and validator/FSM performance.
- Efficiency
- Token volume, output per request, cost, context efficiency, request-error and hallucination rates.
Full scoring definition in the methodology → · Sample transcript
Model leaderboard
| # | Model | Score / 10 | Coverage | n | Last seen |
|---|---|---|---|---|---|
| 1 | Xiaomi Mimo v2.5 | 8.46 | 97% | 3 | 2026-07-26 |
| 2 | Google Gemini Flash 3.6 | 7.96 | 98% | 5 | 2026-07-24 |
| 3 | Qwen Plus qwen3.7 | 7.95 | 98% | 2 | 2026-07-22 |
| 4 | Google Gemini Flash Lite 3.1 † | 7.88 | 62% | 6 | 2026-07-16 |
| 5 | Meta Muse Spark 1.1 | 7.86 | 96% | 2 | 2026-07-19 |
| 6 | Stepfun Step Flash 3.7 † | 7.80 | 96% | 1 | 2026-07-19 |
| 7 | Thinkingmachines Inkling † | 7.31 | 100% | 1 | 2026-07-18 |
| 8 | DeepSeek Deepseek Flash v4 † | 7.12 | 38% | 2 | 2026-06-27 |
| 9 | xAI Grok 4.5 † | 6.99 | 50% | 1 | 2026-07-10 |
| 10 | OpenAI Gpt Terra 5.6 † | 6.89 | 50% | 1 | 2026-07-12 |
| 11 | Google Gemini Flash 3 (preview) † | 5.99 | 48% | 5 | 2026-07-10 |
| 12 | Anthropic Claude Sonnet 5 | 5.80 | 65% | 3 | 2026-07-23 |
| 13 | Google Gemini Pro 3.1 (preview) † | 5.21 | 50% | 1 | 2026-07-03 |
| 14 | Google Gemini Flash 3.5 † | 4.92 | 50% | 1 | 2026-07-05 |
| 15 | Xiaomi Mimo Pro v2.5 † | 4.24 | 38% | 1 | 2026-06-29 |
† low confidence — sparse role or low data coverage (renormalized, never zero-filled).
Per-run history (newest first)
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.