A Multimodal Vision-Language Framework for Financial Anomaly Detection
Abstract
Financial markets are complex systems where state estimation typically relies on either quantitative time-series models or subjective technical analysis. A gap remains in bridging scalable quantitative methods with the pattern-recognition intuition of human experts. In this work, we introduce a multimodal vision-language framework for financial anomaly detection that explores reframing market regime classification as a visual recognition task using advanced multimodal AI. We propose a methodology to project high-dimensional market indicators onto a standardized, multi-panel image format. We leverage the zero-shot capability of state-of-the-art Vision Language Models (VLMs), specifically GPT-4.1. We employ a rolling quantile-based approach for programmatic data labelling to reduce subjectivity. Furthermore, to address the need for transparency, our framework extracts structured natural language rationales directly from the model, offering interpretable insights alongside classifications. Our results, validated through out-of-sample testing on the S&P500, suggest that this multimodal prompting approach achieves promising fidelity in detecting systemic anomalies, indicating its potential as a reliable and interpretable framework for financial AI.
© 2026 Siang-Li JHENG, Daniel Traian PELE, Rahul TAK, Ştefan GAMAN, published by Bucharest University of Economic Studies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.