情态动词
变压器
计算机科学
工程类
电气工程
电压
材料科学
高分子化学
作者
R. Arul Jose,R. Santhana Krishnan,K. Penyameen,J. Relin Francis Raj,V. Vinoth Kumar,M. Geetha Priya
标识
DOI:10.1109/icsadl65848.2025.10933041
摘要
This research presents a Transformer-based multi-modal architecture for predicting box office revenue by integrating diverse data sources: text, visuals, and numerical features. The proposed framework leverages RoBERTa for textual data such as movie descriptions, reviews, and scripts, Vision Transformers (ViT) for visual data including posters and trailers, and MLPs for numerical data like budget, genre, and actor popularity. A cross-attention mechanism is employed to fuse multi-modal embeddings, enabling comprehensive analysis and accurate predictions. The model is trained using PyTorch, optimized with AdamW, and enhanced by data augmentation and TPUs for efficiency. Performance is evaluated using metrics like MAE, RMSE, and R2. Deployment on AWS via FastAPI ensures scalable access, with SHAP explainability providing insights into model decisions. This system offers a robust and interpretable solution for predicting box office performance, catering to the evolving needs of the entertainment industry.
科研通智能强力驱动
Strongly Powered by AbleSci AI