Open-weight agent cores · service-desk task map

Which Open-Weight LLM to Run Your Agent

28 open-weight models scored 0–5 as the reasoning core of an agent across 20 common service-desk tasks, tiered L1 routine, L2 escalated, L3 expert. Agentic ability (tool use, planning, policy adherence, recovery) is scored, not raw chat knowledge. Click any cell for evidence.

Best agent core by tier
Best model by task
Model × agentic-task matrix · sorted by model
0
5
L1L2L3routine → expert
confidence: solid=high, faint=low
Rreasoning model