计算机科学
人工智能
自然语言处理
语音识别
音乐信息检索
语言学
情绪分析
动力学(音乐)
钥匙(锁)
计算语言学
作者
Umangi Nigam,Tushar Swanskar,Rohit Kumar Kaliyar,Mohit Agarwal
标识
DOI:10.1109/temsmet65536.2025.11467389
摘要
We are proposing an emotion-conditioned music generation system that is made up of multimodal inputs, waveform-level analysis, and human-centric evaluation. Going beyond prior models that only use symbolic representations or a rather opaque notion of emotion vectors, we instead marry Transformer- and GAN-based models with fine-grained audio analysis from the Librosa toolkit. Different waveform shape profiles for different emotions–steep peaks for “Angry” and smooth shapes for “Calm”-are tied to subjective validation and FAD or CLAP metrics. Experiments against state-of-the-art models have revealed that this approach brings gains in fidelity, alignment, and listener preference. This work thus takes a step toward actualizing affective AI through real-time generation of emotional music for therapy and entertainment.
科研通智能强力驱动
Strongly Powered by AbleSci AI