MotionGPT: Human Motion as a Foreign Language
AI-generated Key Points
- Advancement of pre-trained large language models has been remarkable
- Building a unified model for language and multi-modal data, such as motion, is largely unexplored
- Human motion exhibits semantic coupling similar to human language
- Combining language data with large-scale motion models enhances performance of motion-related tasks through motion-language pre-training
- MotionGPT is a unified and versatile motion-language model designed for multiple motion-relevant tasks
- Discrete vector quantization is used for human motion and transfer 3D motion into motion tokens
- MotionGPT performs language modeling on both motion and text in a unified manner, treating human motion as its own specific language
- Pre-trained with a mixture of motion-language data and fine-tuned on prompt-based question-and-answer tasks using prompt learning techniques
- Extensive experiments show MotionGPT achieves state-of-the-art performances across various motion tasks including text-driven generation, captioning, prediction, and generating intermediate motions between fixed start and end points.
- MotionGPT combines the strengths of pre-trained language models with the unique characteristics of human motion.
- Offers a uniform approach that treats human motion as a foreign language and leverages the powerful generation and transfer abilities of pre-trained language models.
- Presents an innovative solution for bridging the gap between language and motion.
Authors: Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, Tao Chen
Abstract: Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perceived as a form of body language. By fusing language data with large-scale motion models, motion-language pre-training that can enhance the performance of motion-related tasks becomes feasible. Driven by this insight, we propose MotionGPT, a unified, versatile, and user-friendly motion-language model to handle multiple motion-relevant tasks. Specifically, we employ the discrete vector quantization for human motion and transfer 3D motion into motion tokens, similar to the generation process of word tokens. Building upon this "motion vocabulary", we perform language modeling on both motion and text in a unified manner, treating human motion as a specific language. Moreover, inspired by prompt learning, we pre-train MotionGPT with a mixture of motion-language data and fine-tune it on prompt-based question-and-answer tasks. Extensive experiments demonstrate that MotionGPT achieves state-of-the-art performances on multiple motion tasks including text-driven motion generation, motion captioning, motion prediction, and motion in-between.
Ask questions about this paper to our AI assistant
You can also chat with multiple papers at once here.
Assess the quality of the AI-generated content by voting
Score: 0
Why do we need votes?
Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.
The previous summary was created more than a year ago and can be re-run (if necessary) by clicking on the Run button below.
Similar papers summarized with our AI tools
Navigate through even more similar papers through a
tree representationLook for similar papers (in beta version)
By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.
Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.