Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

AI-generated keywords: Video PreTraining (VPT)

AI-generated Key Points

The license of the paper does not allow us to build upon its content and the key points are generated using the paper metadata rather than the full article.

  • Pretraining on noisy, internet-scale datasets is extensively studied for training models with broad, general capabilities for text, images, and other modalities.
  • Publicly available data in sequential decision domains often lacks the labels required to train behavioral priors in the same way.
  • A team of researchers led by Bowen Baker developed a novel approach called Video PreTraining (VPT), which extends the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning.
  • In VPT, agents learn to act by watching online unlabeled videos.
  • With a small amount of labeled data, they could train an inverse dynamics model accurate enough to label a huge unlabeled source of online data from which they could then train a general behavioral prior.
  • The team's models exhibited human-level performance for many tasks and were even able to craft diamond tools in Minecraft - a task that can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.
  • VPT represents an important step forward in developing more efficient and effective methods for training agents in sequential decision domains where labeled data is scarce or nonexistent.
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, Jeff Clune

Abstract: Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data does not contain the labels required to train behavioral priors in the same way. We extend the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning wherein agents learn to act by watching online unlabeled videos. Specifically, we show that with a small amount of labeled data we can train an inverse dynamics model accurate enough to label a huge unlabeled source of online data -- here, online videos of people playing Minecraft -- from which we can then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), we show that this behavioral prior has nontrivial zero-shot capabilities and that it can be fine-tuned, with both imitation learning and reinforcement learning, to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. For many tasks our models exhibit human-level performance, and we are the first to report computer agents that can craft diamond tools, which can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.

Submitted to arXiv on 23 Jun. 2022

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

The license of the paper does not allow us to build upon its content and the AI assistant only knows about the paper metadata rather than the full article.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2206.11795v1

This paper's license doesn't allow us to build upon its content and the summarizing process is here made with the paper's metadata rather than the article.

The technique of pretraining on noisy, internet-scale datasets has been extensively studied for training models with broad, general capabilities for text, images, and other modalities. However, in sequential decision domains such as robotics, video games, and computer use, publicly available data often lacks the labels required to train behavioral priors in the same way. To address this issue, a team of researchers led by Bowen Baker developed a novel approach called Video PreTraining (VPT), which extends the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning. In VPT, agents learn to act by watching online unlabeled videos. The researchers showed that with a small amount of labeled data they could train an inverse dynamics model accurate enough to label a huge unlabeled source of online data - specifically online videos of people playing Minecraft - from which they could then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), the behavioral prior exhibited nontrivial zero-shot capabilities and could be fine-tuned with both imitation learning and reinforcement learning to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. The team's models exhibited human-level performance for many tasks and were even able to craft diamond tools in Minecraft - a task that can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish. This achievement is significant because it demonstrates that VPT can enable agents to learn complex behaviors from raw sensory inputs without explicit supervision or reward signals. Overall, VPT represents an important step forward in developing more efficient and effective methods for training agents in sequential decision domains where labeled data is scarce or nonexistent. The researchers' findings have implications not only for video game AI but also for robotics and other applications where autonomous agents must make decisions based on sensory inputs.
Created on 06 May. 2023

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

The previous summary was created more than a year ago and can be re-run (if necessary) by clicking on the Run button below.

The license of this specific paper does not allow us to build upon its content and the summarizing tools will be run using the paper metadata rather than the full article. However, it still does a good job, and you can also try our tools on papers with more open licenses.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.