Favur Evals
favur.dev

Model · OpenAI

OpenAI Gpt Luna 5.6

1 appearance · meta-score 64.7/100 · last seen 2026-07-14 · scored under rubric v15

Pooled composite across 1 primary run: 61.1 — no interval, a single run proves nothing about spread provisional: n < 3, treat as unsettled

Agent role scores

Per-agent-role scores for OpenAI Gpt Luna 5.6
Agent roleScore / 10
Sprint-Plan Agent7.99
Scout Agent7.82
Test Agent7.40
Orchestrator7.03
Code-Review Agent6.67
Sprint-Review Agent6.34
Develop Agent6.28
Pseudocode Agent6.24
Build Agent5.84
Code Agent5.58
Platform Agent4.95

Run history (primary model)

Runs where OpenAI Gpt Luna 5.6 was the primary model, oldest first
RunSoWDateComposite
Gpt Luna 5.6circlesJul 202661.1

Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.