Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping be-tween audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, resulting in plain/boring motions during...

Celý popis

Uloženo v:

Podrobná bibliografie
Vydáno v:	Proceedings / IEEE International Conference on Computer Vision s. 11273 - 11282
Hlavní autoři:	Li, Jing, Kang, Di, Pei, Wenjie, Zhe, Xuefei, Zhang, Ying, He, Zhenyu, Bao, Linchao
Médium:	Konferenční příspěvek
Jazyk:	angličtina
Vydáno:	IEEE 01.10.2021
Témata:	Action and behavior recognition Bicycles Codes Computer vision Correlation Gestures and body pose Speech coding Three-dimensional displays Training Vision + other modalities
ISSN:	2380-7504
On-line přístup:	Získat plný text
Tagy:	Přidat tag Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!

Buďte první, kdo okomentuje tento záznam!