The advancement of information technology and the rise of generative AI have paved the way for the development of Large Language Models (LLMs) tailored for TCM diagnostics. However, existing LLMs in the field of TCM face challenges in interpretability, limited modality in interaction, and robustness. To address these limitations, we propose MCM, a Multi-Agent Collaborative Multimodal Framework for TCM Diagnosis. This framework enables robust and interpretable multimodal diagnosis through multi-agent collaboration, offering novel methodologies for applying LLMs in the TCM domain. Experimental results demonstrate that the model within the MCM framework improved performance after fine-tuning, with additional capability gains under the MCM framework’s support, effectively addressing the challenges faced by LLMs in TCM, including interpretability, limited data modality, and lack of robustness. The code is open-sourced at: https://github.com/JerryMazeyu/MCM.