Favur Evals
favur.dev

Model · Qwen

Qwen Plus qwen3.7

4 appearances · meta-score 78.5/100 · last seen 2026-07-22 · scored under rubric v15

Pooled composite across 4 primary runs: 68.9 ± 18.5 (95% CI)

Agent role scores

Per-agent-role scores for Qwen Plus qwen3.7
Agent roleScore / 10
Code-Review Agent8.84
Sprint-Review Agent8.81
Architect8.49
Orchestrator8.45
Develop Agent8.15
Sprint-Plan Agent7.63
Build Agent7.59
Director7.43
Pseudocode Agent7.38
Code Agent6.97
Test Agent6.92
Platform Agent6.59
Scout Agent6.30

Run history (primary model)

Runs where Qwen Plus qwen3.7 was the primary model, oldest first
RunSoWDateComposite
Plus qwen3.7circlesJul 202664.3
Plus qwen3.7circlesJul 202679.9
Plus qwen3.7solarSystemJul 202654.7
Plus qwen3.7circlesJul 202676.7

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.