GraphRoute-Transfer: Topology-Generalizable Routing Convergence Optimization via Graph Reinforcement Learning

Authors

  • Yuto Nakamura University of Massachusetts Amherst, USA

DOI:

https://doi.org/10.54097/h0c18694

Keywords:

Routing convergence, Graph neural networks, Reinforcement learning, OSPF, Timer optimization, Topology generalization, Network reliability

Abstract

Fast and stable routing convergence is critical in large IP networks, and the interior-gateway-protocol (IGP) timers that govern failure detection (Hello/Dead intervals) expose a fundamental tension: aggressive timers detect failures quickly but inflate control overhead and trigger route flaps, whereas conservative timers are stable but slow. Recent work such as DRL-Adapt has shown that deep reinforcement learning can tune these timers better than static defaults, but it operates on a flat, globally-aggregated network state and emits a single network-wide timer, so it can neither exploit the spatial heterogeneity of real topologies nor transfer architecturally across networks of different size. We propose GraphRoute-Transfer, a graph-neural-network policy that assigns per-node timers from local structural features and is by construction permutation- and size-invariant. Because control-plane fragility and failure criticality are spatially heterogeneous, the cost-minimizing timer assignment varies across the graph; our policy learns this mapping and applies it zero-shot to unseen topologies of arbitrary size. Training is guided by a coordinate-descent search oracle on a convergence-cost objective, so the expensive per-topology optimization is amortized into a sub-millisecond inference. On 231 real topologies from the Internet Topology Zoo, GraphRoute-Transfer reduces mean convergence time by 37.3% relative to the OSPF default and to a flat DRL baseline, lowers the composite convergence-cost objective by 16.4% over the flat baseline, and attains 1.317 cost—within 0.3% of the search oracle—while running about 8,160× faster than the search. Crucially, a policy trained only on networks with ≤70 nodes maintains its gains on unseen networks up to 140 nodes, whereas the flat baseline degenerates to a global constant that cannot adapt.

Downloads

Download data is not yet available.

References

[1] Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484-489. https://doi.org/10.1038/nature16961 DOI: https://doi.org/10.1038/nature16961

[2] Knight, S., Nguyen, H. X., Falkner, N., Bowden, R., & Roughan, M. (2011). The Internet topology zoo. IEEE Journal on Selected Areas in Communications, 29(9), 1765-1775. https://doi.org/10.1109/JSAC.2011.111002 DOI: https://doi.org/10.1109/JSAC.2011.111002

[3] Moy, J. (1998). OSPF version 2 (RFC 2328). IETF. DOI: https://doi.org/10.17487/rfc2328

[4] Rekhter, Y., Li, T., & Hares, S. (2006). A Border Gateway Protocol 4 (BGP-4) (RFC 4271). IETF. DOI: https://doi.org/10.17487/rfc4271

[5] Labovitz, C., Ahuja, A., Bose, A., & Jahanian, F. (2001). Delayed Internet routing convergence. IEEE/ACM Transactions on Networking, 9(3), 293-306. https://doi.org/10.1109/90.929852 DOI: https://doi.org/10.1109/90.929852

[6] Francois, P., Filsfils, C., Evans, J., & Bonaventure, O. (2005). Achieving sub-second IGP convergence in large IP networks. ACM SIGCOMM Computer Communication Review, 35(3), 35-44. https://doi.org/10.1145/1070873.1070880 DOI: https://doi.org/10.1145/1070873.1070877

[7] Basu, A., & Riecke, J. G. (2001). Stability issues in OSPF routing. In Proceedings of the ACM SIGCOMM Conference (pp. 225-236). ACM. https://doi.org/10.1145/383059.383075 DOI: https://doi.org/10.1145/383059.383077

[8] Griffin, T. G., & Premore, B. J. (2001). An experimental analysis of BGP convergence time. In Proceedings of the IEEE International Conference on Network Protocols (ICNP) (pp. 53-61). IEEE. https://doi.org/10.1109/ICNP.2001.992887 DOI: https://doi.org/10.1109/ICNP.2001.992760

[9] Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529-533. https://doi.org/10.1038/nature14236 DOI: https://doi.org/10.1038/nature14236

[10] Wang, B., Wang, Z., Zhao, W., Zhang, F., & Shang, W. (2026). DRL-Adapt: Deep reinforcement learning for adaptive routing convergence optimization in large-scale networks. IEEE Open Journal of the Computer Society, 7, 1-14. https://doi.org/10.1109/OJCS.2026.3687441 DOI: https://doi.org/10.1109/OJCS.2026.3687441

[11] van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with double Q-learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 30, No. 1, pp. 2094-2100). AAAI Press. DOI: https://doi.org/10.1609/aaai.v30i1.10295

[12] Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., & de Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML) (pp. 1995-2003). PMLR.

[13] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2016). Prioritized experience replay. In Proceedings of the International Conference on Learning Representations (ICLR).

[14] Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML) (pp. 1928-1937). PMLR.

[15] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv, arXiv:1707.06347. https://doi.org/10.48550/arXiv.1707.06347

[16] Sutton, R. S., McAllester, D., Singh, S., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems 13 (NIPS) (pp. 1057-1063).

[17] Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., & Dahl, G. E. (2017). Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning (ICML) (pp. 1263-1272). PMLR.

[18] Kipf, T. N., & Welling, M. (2017). Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations (ICLR).

[19] Hamilton, W. L., Ying, R., & Leskovec, J. (2017). Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems 30 (NeurIPS) (pp. 1024-1034).

[20] Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., & Bengio, Y. (2018). Graph attention networks. In Proceedings of the International Conference on Learning Representations (ICLR).

[21] Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., & Monfardini, G. (2009). The graph neural network model. IEEE Transactions on Neural Networks, 20(1), 61-80. https://doi.org/10.1109/TNN.2008.2005605 DOI: https://doi.org/10.1109/TNN.2008.2005605

[22] Bengio, Y., Lodi, A., & Prouvost, A. (2021). Machine learning for combinatorial optimization: A methodological tour d'horizon. European Journal of Operational Research, 290(2), 405-421. https://doi.org/10.1016/j.ejor.2020.07.063 DOI: https://doi.org/10.1016/j.ejor.2020.07.063

[23] Teng, D., Rhee, M., Qin, Y., Zi, B., & Liu, W. (2026). SW-SpeedDLM: Sliding-window speculative decoding for diffusion language models under long-context constraints. Mathematics, 14(12), Article 2137. https://doi.org/10.3390/math14122137 DOI: https://doi.org/10.3390/math14122137

[24] Wang, Z., Yang, J. S., Shang, W., & Ding, J. (2026). FairPromote: Explainable and fairness-aware talent promotion prediction via adversarial debiasing and SHAP-based interpretation. IEEE Access, 14, 1-16. DOI: https://doi.org/10.1109/ACCESS.2026.3692578

[25] Teng, D. (2025). TEAS: Token- and energy-aware autoscaling for cost-efficient LLM serving. AI and Data Science Journal, 6(3), 1-10.

[26] Zhang, F., Guo, Z., Ding, J., Yang, J., & Liu, W. (2026). Adaptive sensor fusion for robust perception in dense fog: A gated vision and LiDAR integration framework. Sensors, 26(12), Article 3728. https://doi.org/10.3390/s26123728 DOI: https://doi.org/10.3390/s26123728

[27] Zi, B. (2024). Large language models for enterprise workflow automation in financial operations. Innovation and Technology Studies, 1(1), 24-29. DOI: https://doi.org/10.61784/its3023

[28] Ding, J., Shen, Z., & Liu, W. (2026). Game-theoretic cost-sensitive adversarial training for robust cloud intrusion detection against GAN-based evasion attacks. Applied Sciences, 16(8), Article 3944. https://doi.org/10.3390/app16083944 DOI: https://doi.org/10.3390/app16083944

[29] Zi, B. (2024). Cloud-native distributed systems for real-time payment intelligence. AI and Data Science Journal, 1(1), 51-56. DOI: https://doi.org/10.61784/adsj3035

[30] Liang, Y., Jiao, Y., Ping, W., Fan, H., & Han, X. (2026). Adaptive event-driven labeling: A neuro-symbolic multiagent framework for causal inference in non-stationary time series. IEEE Access, 14, 1-18. DOI: https://doi.org/10.1109/ACCESS.2026.3709267

Downloads

Published

20-07-2026

Issue

Section

Articles

How to Cite

Nakamura, Y. (2026). GraphRoute-Transfer: Topology-Generalizable Routing Convergence Optimization via Graph Reinforcement Learning. Journal of Computing and Electronic Information Management, 22(1), 12-18. https://doi.org/10.54097/h0c18694