Skip to content Text size 100%
Canon · Agnostic storm-review

The agent scoreboard — quality versus cost, by kind of task

What it is

Every review-panel run is recorded: which agents took part, which model graded, each agent's score, tokens and exact cost. A leaderboard then shows quality versus cost per agent, per kind of task. Over time the system learns which model to trust for which job, and what that trust costs.

The data

  • Panel runs — the kind of task, a short prompt excerpt, the grader, the number of members and the consensus pick, with tenant scoping on every row.
  • Panel scores — per agent: provider, model, anonymous label, a 0–10 score with the grader's reason, tokens, cost in dollars, and whether it was the consensus pick.
  • Both use row-level security, closed by default.

The metrics

  • Quality = the grader's score (an AI judge, so it's noisy; the count of runs is always shown).
  • Cost = the provider's exact reported usage cost.
  • Efficiency = score per dollar. Free local models are shown separately, never as infinite.
  • Win rate = how often an agent is picked as the consensus answer.

Keeping it honest

  • Graders are anonymised and rotated. Scores are aggregated across graders.
  • Small samples are greyed out rather than ranked.
  • The primary metric and stopping rule are decided in advance, and the study is ready to report a null result.