Model · OpenAI
OpenAI Gpt Luna 5.6
1 appearance · meta-score 64.7/100 · last seen 2026-07-14 · scored under rubric v15
Pooled composite across 1 primary run: 61.1 — no interval, a single run proves nothing about spread provisional: n < 3, treat as unsettled
- $26.22Total cost
- 83,035,988Tokens
- 82.2%Cache hit
- 1,568Requests
- 0.00%Failure rate
- 1.14Tools/response
- 3,166,753Tokens per $
Agent role scores
| Agent role | Score / 10 |
|---|---|
| Sprint-Plan Agent | 7.99 |
| Scout Agent | 7.82 |
| Test Agent | 7.40 |
| Orchestrator | 7.03 |
| Code-Review Agent | 6.67 |
| Sprint-Review Agent | 6.34 |
| Develop Agent | 6.28 |
| Pseudocode Agent | 6.24 |
| Build Agent | 5.84 |
| Code Agent | 5.58 |
| Platform Agent | 4.95 |
Run history (primary model)
| Run | SoW | Date | Composite |
|---|---|---|---|
| Gpt Luna 5.6 | circles | Jul 2026 | 61.1 |
Every run on Favur Evals is scored by the same deterministic engine, on the same Statements of Work. The benchmark is self-funded — no vendor sponsorship, credits, or grants.