DriftGuard-TriAudit: Concept-Drift-Aware Continual Multimodal Learning for Evolving Financial Statement Fraud Detection
DOI:
https://doi.org/10.54097/5bjxgq31Keywords:
Financial statement fraud detection, Concept drift, Continual learning, Multimodal learning, Audit evidence chain, explainable artificial intelligenceAbstract
Financial statement fraud detection is usually trained as a static supervised classification task, yet real reporting behavior evolves as firms change disclosure style, enforcement priorities shift, and fraudsters adapt to known audit screens. This paper proposes DriftGuard-TriAudit, a concept-drift-aware continual multimodal learning framework for evolving financial statement fraud detection. Building positively on the recent TriAudit design, which demonstrates the value of jointly reasoning over financial ratios, management discussion text, and document-layout evidence, the proposed framework adds three mechanisms for non-stationary audit settings: modality-wise drift diagnostics, reliability-aware gated fusion, and replay-regularized continual updating. The model monitors population stability and calibration performance in numerical, textual, and visual-layout streams, then adapts the contribution of each modality while preserving prior fraud evidence through a compact replay memory. Because a full EDGAR-plus-AAER real-data construction was not available in the execution environment, the empirical section reports only a reproducible synthetic benchmark that explicitly simulates four temporal fraud regimes and does not claim to be real company data. Across five seeds, DriftGuard-TriAudit obtains the strongest average precision (0.628±0.087) and F1 score (0.579±0.061) among evaluated practical models, with substantially higher precision than naive continual fine-tuning (0.547 vs. 0.428) while maintaining comparable AUC-ROC (0.915 vs. 0.924). The results indicate that drift-aware multimodal fusion can reduce false positives under evolving fraud mechanisms, and the released code provides a direct path for replacing the synthetic generator with EDGAR/AAER features.
Downloads
References
[1] Beneish, M. D. (1999). The detection of earnings manipulation. Financial Analysts Journal, 55(5), 24-36. https://doi.org/10.2469/faj.v55.n5.2296 DOI: https://doi.org/10.2469/faj.v55.n5.2296
[2] Ping, W., Jiao, Y., Fan, H., & Zhang, X. (2026). Multimodal fraud detection in financial statements: A trimodal attention network with contrastive evidence chain construction. IEEE Access, 14, 80456-80468. https://doi.org/10.1109/ACCESS.2026.3695466 DOI: https://doi.org/10.1109/ACCESS.2026.3695466
[3] Dechow, P. M., Ge, W., Larson, C. R., & Sloan, R. G. (2011). Predicting material accounting misstatements. Contemporary Accounting Research, 28(1), 17-82. https://doi.org/10.1111/j.1911-3846.2010.01041.x DOI: https://doi.org/10.1111/j.1911-3846.2010.01041.x
[4] Perols, J. L. (2011). Financial statement fraud detection: An analysis of statistical and machine learning algorithms. Auditing: A Journal of Practice & Theory, 30(2), 19-50. https://doi.org/10.2308/ajpt-50009 DOI: https://doi.org/10.2308/ajpt-50009
[5] Cecchini, M., Aytug, H., Koehler, G. J., & Pathak, P. (2010). Detecting management fraud in public companies. Management Science, 56(7), 1146-1160. https://doi.org/10.1287/mnsc.1100.1174 DOI: https://doi.org/10.1287/mnsc.1100.1174
[6] Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. The Journal of Finance, 66(1), 35-65. https://doi.org/10.1111/j.1540-6261.2010.01625.x DOI: https://doi.org/10.1111/j.1540-6261.2010.01625.x
[7] Li, F. (2008). Annual report readability, current earnings, and earnings persistence. Journal of Accounting and Economics, 45(2-3), 221-247. https://doi.org/10.1016/j.jacceco.2008.02.003 DOI: https://doi.org/10.1016/j.jacceco.2008.02.003
[8] Loukas, L., Fergadiotis, M., Androutsopoulos, I., & Malakasiotis, P. (2021). EDGAR-CORPUS: Billions of tokens make the world go round. In Proceedings of the 4th Workshop on Economics and Natural Language Processing (ECONLP). arXiv, arXiv:2109.14394. https://doi.org/10.48550/arXiv.2109.14394 DOI: https://doi.org/10.18653/v1/2021.econlp-1.2
[9] U.S. Securities and Exchange Commission. (n.d.). Accounting and auditing enforcement releases. SEC official collection. https://www.sec.gov/enforce/accounting-and-auditing-enforcement-releases
[10] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), Article 44. https://doi.org/10.1145/2523813 DOI: https://doi.org/10.1145/2523813
[11] Gama, J., Medas, P., Castillo, G., & Rodrigues, P. (2004). Learning with drift detection. In A. L. C. Bazzan & S. Labidi (Eds.), Advances in Artificial Intelligence: SBIA 2004 (pp. 286-295). Springer. https://doi.org/10.1007/978-3-540-28645-5_29 DOI: https://doi.org/10.1007/978-3-540-28645-5_29
[12] Bifet, A., & Gavaldà, R. (2007). Learning from time-changing data with adaptive windowing. In Proceedings of the SIAM International Conference on Data Mining (SDM) (pp. 443-448). SIAM. https://doi.org/10.1137/1.9781611972771.42 DOI: https://doi.org/10.1137/1.9781611972771.42
[13] Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., & Hadsell, R. (2017). Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the United States of America, 114(13), 3521-3526. https://doi.org/10.1073/pnas.1611835114 DOI: https://doi.org/10.1073/pnas.1611835114
[14] Rebuffi, S.-A., Kolesnikov, A., Sperl, G., & Lampert, C. H. (2017). iCaRL: Incremental classifier and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5533-5542). IEEE. https://doi.org/10.1109/CVPR.2017.587 DOI: https://doi.org/10.1109/CVPR.2017.587
[15] Lopez-Paz, D., & Ranzato, M. (2017). Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems 30 (NeurIPS).
[16] Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T., & Wayne, G. (2019). Experience replay for continual learning. In Advances in Neural Information Processing Systems 32 (NeurIPS).
[17] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (NeurIPS).
[18] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Volume 1 (Long and Short Papers) (pp. 4171-4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423 DOI: https://doi.org/10.18653/v1/N19-1423
[19] Araci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language models. arXiv, arXiv:1908.10063. https://doi.org/10.48550/arXiv.1908.10063
[20] Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., & Zhou, M. (2021). LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP) (pp. 2579-2591). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.201 DOI: https://doi.org/10.18653/v1/2021.acl-long.201
[21] van den Oord, A., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv, arXiv:1807.03748. https://doi.org/10.48550/arXiv.1807.03748
[22] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning (ICML) (pp. 1597-1607). PMLR.
[23] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD) (pp. 785-794). ACM. https://doi.org/10.1145/2939672.2939785 DOI: https://doi.org/10.1145/2939672.2939785
[24] Wang, Z., Yang, J. S., Shang, W., & Ding, J. (2026). FairPromote: Explainable and fairness-aware talent promotion prediction via adversarial debiasing and SHAP-based interpretation. IEEE Access, 14, 1-16. DOI: https://doi.org/10.1109/ACCESS.2026.3692578
[25] Teng, D. (2025). TEAS: Token- and energy-aware autoscaling for cost-efficient LLM serving. AI and Data Science Journal, 6(3), 1-10.
[26] Zhang, F., Guo, Z., Ding, J., Yang, J., & Liu, W. (2026). Adaptive sensor fusion for robust perception in dense fog: A gated vision and LiDAR integration framework. Sensors, 26(12), Article 3728. https://doi.org/10.3390/s26123728 DOI: https://doi.org/10.3390/s26123728
[27] Zi, B. (2024). Large language models for enterprise workflow automation in financial operations. Innovation and Technology Studies, 1(1), 24-29. DOI: https://doi.org/10.61784/its3023
[28] Ding, J., Shen, Z., & Liu, W. (2026). Game-theoretic cost-sensitive adversarial training for robust cloud intrusion detection against GAN-based evasion attacks. Applied Sciences, 16(8), Article 3944. https://doi.org/10.3390/app16083944 DOI: https://doi.org/10.3390/app16083944
[29] Zi, B. (2024). Cloud-native distributed systems for real-time payment intelligence. AI and Data Science Journal, 1(1), 51-56. DOI: https://doi.org/10.61784/adsj3035
[30] Teng, D. (2025). PACO: Predictive auto-configuration for SLO-constrained large language model inference serving. Innovation and Technology Studies, 2(1), 1-10.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computing and Electronic Information Management

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








