Large help desks route hundreds of incidents a day to dozens of resolver groups, and a wrong assignment means bounce-backs, delays and broken SLAs. The dominant industry approach treats dispatch as text classification — but a classifier that learns “which group did this ticket go to” ignores the coupling between decisions. A reinforcement-learning agent can learn to divert tickets when the preferred queue is congested.

We formulate incident routing as a Markov decision process and solve it with tabular methods, linear approximation and deep reinforcement learning (DQN, Double DQN and actor–critic A2C). The simulation environment is calibrated end to end against 177,940 real incidents from a tax authority’s help desk (2021–2026): arrivals as a non-homogeneous Poisson process simulated by thinning, service times drawn from empirical per-group distributions, steady-state warm-up fixed by Little’s law, and SLA targets set at the historical 75th percentile per ticket type.

With a state space of 8.3 × 10¹⁵ configurations the Q-table collapses — roughly 0.01% coverage in evaluation — which is what motivates function approximation. Under a multi-replica protocol with 95% confidence intervals and Holm–Bonferroni correction, the linear approximator (46→15) beats the historical dispatch heuristic with statistical significance, while the deep DQN (46→128→128→15) falls below it. The cause is maximization bias, measured directly as the gap between the predicted Q value and the return actually obtained; Double DQN reduces that bias significantly (p = 0.016), confirming the mechanism without reversing the deficit. The conclusion: on problems with moderate signal and heavy-tailed noise, the advantage comes from the network’s ability to generalize rather than from its depth — and a replicated statistical protocol is essential if you don’t want to draw spurious conclusions.

Your browser can’t display the PDF inline. Open PDF

← All research