28 open-weight models scored 0–5 as the reasoning core of an agent across 20 common service-desk tasks, tiered L1 routine, L2 escalated, L3 expert. Agentic ability (tool use, planning, policy adherence, recovery) is scored, not raw chat knowledge. Click any cell for evidence.