运动(物理)
计算机科学
代表(政治)
集合(抽象数据类型)
人工智能
代码段
采样(信号处理)
计算机视觉
空格(标点符号)
文本生成
功能(生物学)
情报检索
程序设计语言
操作系统
法学
滤波器(信号处理)
政治
生物
进化生物学
政治学
作者
Chuan Guo,Shihao Zou,Xinxin Zuo,Sen Wang,Wei Ji,Xingyu Li,Li Cheng
出处
期刊:
日期:2022-06-01
卷期号:: 5142-5151
被引量:392
标识
DOI:10.1109/cvpr52688.2022.00509
摘要
Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions.
科研通智能强力驱动
Strongly Powered by AbleSci AI