摘要
Clinical decision support systems now provide more options because of the integration of medical imaging, clinical text, and contextual data utilizing contemporary visionlanguage models (VLMs). Performance, explainability, accessibility, and regulatory compliance are all trade-offs in any model. In this study, we assess four VLMs using multimodal medical benchmarks: MedGemma-4B-IT, LLaVA-Med, MedVLMR1, and GPT-4V. Because of its high multimodal capabilities, reasoning ability, integration of external knowledge, and modifiability, we have chosen MedGemma-4B-IT as the best opensource choice. MedGemma-4B-IT provides competitive baseline performance on standard medical VQA benchmarks by building upon the Gemma-3 architecture ($\mathbf{4 B}$ parameters) with a MedSigLIP visual encoder. We built a modular pipeline around this model to enable practical application, including explainable AI approaches, data validation, automated text de-identification, external knowledge retrieval, and integration of vital signs generated from the Internet of Things. We can customize outputs for the end user, which may be a patient or a clinician, thanks to the system’s explainability characteristics. This method still has drawbacks, such as single-image input limitations and the absence of official clinical confirmation. However, our findings show that MedGemma-4B-IT can be effectively integrated into an explainable, IoT-augmented multimodal pipeline, representing a significant step toward trustworthy AI systems in healthcare.