Model × Scenario Matrix
One column per model, every model against every scenario, scored modal-of-5 on the commit-permission gate decision. Each square is split: the left half is the model at its reasoning floor (off / minimal / low), the right half is the same model at high reasoning (or "on" for DeepSeek's binary on/off), so the effect of reasoning is visible per scenario. Green = trials right, red = wrong, darker = more unanimous. Red means wrong relative to the row's right call: on must-proceed rows it is over-refusal, and on must-hold rows it is under-refusal. Rows sort hardest first, so the top band shows where the gate fails across families. Hover any square for the full per-model breakdown. The footer row is per-run accuracy (floor│high), the leaderboard read down the columns.
Loading matrix…
Loading matrix…