Skip to main navigation Skip to search Skip to main content

Towards Open Domain Text-Driven Synthesis of Multi-person Motions

  • Mengyi Shan
  • , Lu Dong
  • , Yutao Han
  • , Yuan Yao
  • , Tao Liu
  • , Ifeoma Nwogu
  • , Guo Jun Qi
  • , Mitch Hill
  • University of Washington
  • SUNY Buffalo
  • Innopeak Technology
  • University of Rochester
  • Westlake University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages67-86
Number of pages20
ISBN (Print)9783031736490
DOIs
StatePublished - 2025
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: Sep 29 2024Oct 4 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15123 LNCS

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period09/29/2410/4/24

Keywords

  • Human Motion Generation
  • Human Pose Dataset
  • Multi-Person Motion Generation
  • Text-to-Motion Generation

Fingerprint

Dive into the research topics of 'Towards Open Domain Text-Driven Synthesis of Multi-person Motions'. Together they form a unique fingerprint.

Cite this