隐藏字幕
计算机科学
语音识别
循环神经网络
视听
内容(测量理论)
多媒体
人工智能
图像(数学)
人工神经网络
数学
数学分析
作者
Agus Hermanto,Giat Karyono,Imam Tahyudin,Boby Sandityas Prahasto
标识
DOI:10.35970/jinita.v7i1.2788
摘要
The primary objective of this research is to develop an image captioning and audio conversion system based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) with the integration of an Attention Mechanism, aimed at improving accessibility for visually impaired individuals. The research design follows a systematic approach involving data collection, preprocessing, model development, training, evaluation, and implementation. The methodology utilizes CNN for visual feature extraction, RNN for language modeling, and an Attention Mechanism to enhance contextual relevance in caption generation. Google Text-to-Speech (gTTS) is also integrated to convert generated captions into audio format. The main outcomes demonstrate that the model is capable of generating coherent and contextually relevant captions, as validated through qualitative assessment and quantitative measurement using the BLEU score. Experimental results show decreasing training and validation loss over 8 epochs without signs of overfitting, indicating stable model performance. The attention visualization confirms the model’s ability to focus on relevant image regions during caption generation. In conclusion, the proposed CNN-RNN architecture with Attention effectively generates descriptive captions and converts them into speech, showing strong potential for real-world accessibility applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI