Favur Evals
favur.dev

Agent role · weight 6/100 · rubric v15

🎬 Director

Read-only strategic planner — defines sprint-sized goals with binary acceptance criteria and writes the phase direction document.

Provisional: n < 15 — the confidence interval is wide; treat rankings as unsettled.

What it's scored on

Output quality
Direction quality: goal presence plus open-question closure ratio (resolved or deferred vs found) and validator/FSM performance.
Efficiency
Token volume, output per request, cost, context efficiency, request-error and hallucination rates.

Model leaderboard

Every model ranked by Director score
#ModelScore / 10CoveragenLast seen
1DeepSeek Deepseek Flash v48.6068%12026-06-26
2Qwen Plus qwen3.77.4396%12026-07-19
3Google Gemini Flash 3.67.4196%12026-07-22
4Meta Muse Spark 1.16.3296%22026-07-19
5Stepfun Step Flash 3.74.3696%12026-07-19

† low confidence — sparse role or low data coverage (renormalized, never zero-filled).

Per-run history (newest first)

Every appearance of the Director role, newest first
DateRunSoWModelScore / 10
2026-07-22the2048__gemini-multi-model-7-22-2026the2048Google Gemini Flash 3.67.41
2026-07-19solarSystem__qwen-qwen3.7-plus-7.18.2026solarSystemQwen Plus qwen3.77.43
2026-07-19solarSystem__stepfun-step-3.7-flash-7.19.2026solarSystemStepfun Step Flash 3.74.36
2026-07-19the2048__meta-muse-spark-1.1-7.19.2026the2048Meta Muse Spark 1.14.94
2026-07-17solarSystem__meta-muse-spark-1.1-7.17.2026solarSystemMeta Muse Spark 1.17.69
2026-06-26the2048__deepseek-deepseek-v4-flash-6.26.2026the2048DeepSeek Deepseek Flash v48.60

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.