Favur Evals
favur.dev

Model · DeepSeek

DeepSeek Deepseek Pro v4

3 appearances · meta-score 58.6/100 · last seen 2026-07-05 · scored under rubric v15

Pooled composite across 3 primary runs: 53.9 ± 5.0 (95% CI)

Agent role scores

Per-agent-role scores for DeepSeek Deepseek Pro v4
Agent roleScore / 10
Sprint-Plan Agent7.58
Scout Agent7.43
Sprint-Review Agent6.78
Spec Agent6.60
Test Agent6.22
Develop Agent6.14
Code-Review Agent5.56
Pseudocode Agent5.25
Platform Agent5.19
Build Agent4.89
Code Agent4.82
Orchestrator4.68

Run history (primary model)

Runs where DeepSeek Deepseek Pro v4 was the primary model, oldest first
RunSoWDateComposite
Deepseek Pro v4circlesJun 202653.9
Deepseek Pro v4circlesJul 202651.9
Deepseek Pro v4circlesJul 202655.9

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.