Transformer Architectures for Multivariate Time Series Forecasting
DOI:
https://doi.org/10.54097/8d6vbb11Keywords:
Multivariate time series, Deep learning, Transformer, Attention mechanism, Feature extraction, Foundation modelsAbstract
Multivariate time series data constitutes the fundamental basis for decision making across real-world scenarios such as financial trading, intensive care, energy dispatch and the Internet of Things. Traditional statistical methods and early deep learning architectures including recurrent neural networks and convolutional neural networks expose structural bottlenecks when these models capture long-range dynamic dependencies, process high-dimensional cross correlations and accommodate data expansion on a massive scale. The Transformer architecture relies on the self-attention mechanism to break the traditional paradigm of sequence modeling through parallel computing capabilities and flexible global receptive fields. This architecture has emerged as the cutting-edge foundation for the analysis of multivariate time series. However, the standard Transformer model remains restricted by quadratic computational complexity, data starvation effects and the lack of interpretability associated with black box characteristics when this model processes real-world data featuring high-frequency fluctuations, non-stationarity and significant noise interference. This paper systematically synthesizes the key theoretical and methodological breakthroughs in the field of time series forecasting in recent years. This work starts from the underlying feature representation mechanism to deeply deconstruct data representation strategies which include dimensionality reduction decomposition, time-frequency domain conversion and masked pre-training. By integrating specific application scenarios, this paper conducts a fine-grained dynamic classification and performance evaluation of current deep learning models. This evaluation specifically focuses on efficient sparse attention, sequence decomposition, patch-based designs and multimodal hybrid Transformer architectures. Building upon this foundation, this work examines the severe challenges facing the field and prospectively explores future academic trajectories which encompass time series foundation models, physics-informed design, counterfactual reasoning and federated learning. The ultimate goal of this paper is to provide a panoramic reference for the theoretical breakthroughs and engineering deployment of complex time series analysis.
Downloads
References
[1] Box, G. E. P., Jenkins, G. M., Reinsel, G. C., & Ljung, G. M. (2015). Time series analysis: Forecasting and control. John Wiley & Sons.
[2] Bai, S., Kolter, J. Z., & Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv, arXiv:1803.01271. https://doi.org/10.48550/arXiv.1803.01271
[3] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (NeurIPS) (pp. 5998-6008).
[4] Lim, B., Arık, S. Ö., Loeff, N., & Pfister, T. (2021). Temporal fusion transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4), 1748-1764. https://doi.org/10.1016/j.ijforecast.2021.03.012
[5] Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., & Zhang, W. (2021). Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 35, No. 12, pp. 11106-11115). AAAI Press. https://doi.org/10.1609/aaai.v35i12.17325
[6] Wu, H., Xu, J., Wang, J., & Long, M. (2021). Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Advances in Neural Information Processing Systems 34 (NeurIPS) (pp. 22419-22430).
[7] Nie, Y., Nguyen, N. H., Sinthong, P., & Kalagnanam, J. (2022). A time series is worth 64 words: Long-term forecasting with transformers. arXiv, arXiv:2211.14730. https://doi.org/10.48550/arXiv.2211.14730
[8] Garza, A., & Mergenthaler-Canseco, M. (2023). TimeGPT-1. arXiv, arXiv:2310.03589. https://doi.org/10.48550/arXiv.2310.03589
[9] Shumway, R. H., & Stoffer, D. S. (2000). Time series analysis and its applications. Springer.
[10] Blázquez-García, A., Conde, A., Mori, U., & Lozano, J. A. (2021). A review on outlier/anomaly detection in time series data. ACM Computing Surveys, 54(3), Article 56. https://doi.org/10.1145/3444690
[11] Liu, P., Lin, J., & Zhang, C. (2024). Heterogeneous multivariate functional time series modeling: A state space approach. IEEE Transactions on Knowledge and Data Engineering, 36(12), 8421-8433. https://doi.org/10.1109/TKDE.2024.3402765
[12] Wang, Z., Fan, J., Wu, H., Sun, D., & Wu, J. (2024). Representing multiview time-series graph structures for multivariate long-term time-series forecasting. IEEE Transactions on Artificial Intelligence, 5(6), 2651-2662. https://doi.org/10.1109/TAI.2023.3332564
[13] Saravana, M. K., Roopa, M. S., Arunalatha, J. S., & Venugopal, K. R. (2026). Transformers for multivariate time series forecasting: Comprehensive analysis, challenges, research opportunities, and future prospects. IEEE Access, 14, 11424-11457.
[14] Xu, H. (2025). AFMDS: Attention-based forecasting model against distribution shift for long-term non-stationary time series prediction. IEEE Access, 13, 155797-155807.
[15] Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. Advances in Neural Information Processing Systems, 27, 3104-3112.
[16] Perslev, M., Jensen, M., Darkner, S., Jennum, P. J., & Igel, C. (2019). U-Time: A fully convolutional network for time series segmentation applied to sleep staging. In Advances in Neural Information Processing Systems 32 (NeurIPS).
[17] Yoon, J., Jarrett, D., & van der Schaar, M. (2019). Time-series generative adversarial networks. In Advances in Neural Information Processing Systems 32 (NeurIPS).
[18] Esteban, C., Hyland, S. L., & Rätsch, G. (2017). Real-valued (medical) time series generation with recurrent conditional GANs. arXiv, arXiv:1706.02633. https://doi.org/10.48550/arXiv.1706.02633
[19] Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., & Jin, R. (2022). FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In Proceedings of the 39th International Conference on Machine Learning (ICML) (pp. 27268-27286). PMLR.
[20] Lu, K., Huo, M., Li, Y., Zhu, Q., & Chen, Z. (2025). CT-PatchTST: Channel-time patch time-series transformer for long-term renewable energy forecasting. arXiv, arXiv:2501.08620. https://doi.org/10.48550/arXiv.2501.08620
[21] Xu, S., Wang, Y., Xu, X., Shi, G., Huang, H., & Zheng, Y. (2025). A PatchTST-GRU based heterogeneous seq2seq model with numerical weather prediction refinement for multi-step wind power forecasting. Scientific Reports, 15(1), Article 16547. https://doi.org/10.1038/s41598-025-98765-4
[22] Wang, Z., Xu, X., Zhang, W., Trajcevski, G., Zhong, T., & Zhou, F. (2022). Learning latent seasonal-trend representations for time series forecasting. In Advances in Neural Information Processing Systems 35 (NeurIPS) (pp. 38775-38787).
[23] Yi, K., Zhang, Q., Fan, W., Wang, S., Wang, P., He, H., Lian, D., An, N., Cao, L., & Niu, Z. (2023). Frequency-domain MLPs are more effective learners in time series forecasting. arXiv, arXiv:2311.06184. https://doi.org/10.48550/arXiv.2311.06184
[24] Zhou, T., Ma, Z., Wen, Q., Sun, L., Yao, T., Yin, W., & Jin, R. (2022). FiLM: Frequency improved Legendre memory model for long-term time series forecasting. In Advances in Neural Information Processing Systems 35 (NeurIPS) (pp. 12677-12690).
[25] Zhang, W., Yang, L., Geng, S., & Hong, S. (2022). Self-supervised time series representation learning via cross reconstruction transformer. arXiv, arXiv:2205.09928. https://doi.org/10.48550/arXiv.2205.09928
[26] Shao, Z., Zhang, Z., Wang, F., & Xu, Y. (2022). Pre-training enhanced spatial-temporal graph neural network for multivariate time series forecasting. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) (pp. 1567-1577). ACM. https://doi.org/10.1145/3534678.3539396
[27] Li, Z., Rao, Z., Pan, L., Wang, P., & Xu, Z. (2023). Ti-MAE: Self-supervised masked time series autoencoders. arXiv, arXiv:2301.08871. https://doi.org/10.48550/arXiv.2301.08871
[28] Zhao, J., Chu, F., Xie, L., Che, Y., Wu, Y., & Burke, A. F. (2026). A survey of transformer networks for time series forecasting. Computer Science Review, 60, Article 100883. https://doi.org/10.1016/j.cosrev.2026.100883
[29] Sun, W., Liu, Z., Yuan, C., Zhou, X., Pei, Y., & Wei, C. (2025). RCSAN: Residual enhanced channel spatial attention network for stock price forecasting. Scientific Reports, 15(1), Article 21800. https://doi.org/10.1038/s41598-025-87654-3
[30] Song, Z., Lu, Q., Xu, H., Zhu, H., Buckeridge, D., & Li, Y. (2024). TimelyGPT: Extrapolatable transformer pre-training for long-term time-series forecasting in healthcare. In Proceedings of the 15th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics (pp. 1-10). ACM. https://doi.org/10.1145/3699767.3699791
[31] Kong, X., Chen, Z., Liu, W., Ning, K., Zhang, L., Marier, S. M., Liu, Y., Chen, Y., & Xia, F. (2025). Deep learning for time series forecasting: A survey. International Journal of Machine Learning and Cybernetics, 16, 5079-5112. https://doi.org/10.1007/s13042-025-02345-6
[32] Chen, C., Petty, K., Skabardonis, A., Varaiya, P., & Jia, Z. (2001). Freeway performance measurement system: Mining loop detector data. Transportation Research Record, 1748(1), 96-102. https://doi.org/10.3141/1748-12
[33] Makridakis, S., Spiliotis, E., & Assimakopoulos, V. (2018). The M4 competition: Results, findings, conclusion and way forward. International Journal of Forecasting, 34(4), 802-808. https://doi.org/10.1016/j.ijforecast.2018.06.001
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computing and Electronic Information Management

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








