Towards Language-Driven Video Inpainting via Multimodal Large Language Models
Abstract: We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that depend on manually labeled binary masks, a process often tedious and labor-intensive. We present the Remove Objects from Videos by Instructions (ROVI) dataset, containing 5,650 videos and 9,091 inpainting results, to support training and evaluation for this task. We also propose a novel diffusion-based language-driven video inpainting framework, the first end-to-end baseline for this task, integrating Multimodal LLMs to understand and execute complex language-based inpainting requests effectively. Our comprehensive results showcase the dataset's versatility and the model's effectiveness in various language-instructed inpainting scenarios. We will make datasets, code, and models publicly available.
- Instructpix2pix: Learning to follow image editing instructions. In CVPR, 2023.
- Language models are few-shot learners. NeurIPS, 2020.
- Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
- Free-form video inpainting with 3d gated convolution and temporal patchgan. In ICCV, 2019.
- An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
- Video inpainting with short-term windows: application to object removal and error concealment. TIP, 2015.
- Actor and action video segmentation from a sentence. In CVPR, 2018.
- Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041, 2023.
- Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022.
- Denoising diffusion probabilistic models. NeurIPS, 2020.
- Proposal-based video completion. In ECCV, 2020.
- Gqa: A new dataset for real-world visual reasoning and compositional question answering. CVPR, 2019.
- Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023.
- Segment anything. arXiv preprint arXiv:2304.02643, 2023.
- Learning blind video temporal consistency. In ECCV, 2018.
- Copy-and-paste networks for deep video inpainting. In ICCV, 2019.
- Short-term and long-term context aggregation network for video inpainting. In ECCV, 2020a.
- Mimic-it: Multi-modal in-context instruction tuning. arXiv preprint arXiv:2306.05425, 2023a.
- Recurrent feature reasoning for image inpainting. In CVPR, 2020b.
- MAT: Mask-aware transformer for large hole image inpainting. In CVPR, 2022a.
- Transformer-based visual segmentation: A survey. arXiv pre-print, 2023b.
- Towards an end-to-end framework for flow-guided video inpainting. In CVPR, 2022b.
- Image inpainting for irregular holes using partial convolutions. In ECCV, 2018.
- Partial convolution for padding, inpainting, and image synthesis. TPAMI, 2022.
- Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023a.
- Fuseformer: Fusing fine-grained information in transformers for video inpainting. In ICCV, 2021a.
- Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023b.
- Deep learning face attributes in the wild. In ICCV, 2015.
- Swin Transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021b.
- Repaint: Inpainting using denoising diffusion probabilistic models. In CVPR, 2022.
- Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021.
- AdaViT: Adaptive vision transformers for efficient image recognition. In CVPR, 2022.
- Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022.
- OpenAI. Gpt-4 technical report, 2023.
- Context encoders: Feature learning by inpainting. In CVPR, 2016.
- A benchmark dataset and evaluation methodology for video object segmentation. In CVPR, 2016.
- Detgpt: Detect what you need via reasoning. arXiv preprint arXiv:2305.14167, 2023.
- Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
- DynamicViT: Efficient vision transformers with dynamic token sparsification. In NeurIPS, 2021.
- High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
- Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 2022.
- Urvos: Unified referring video object segmentation network with a large-scale benchmark. In ECCV, 2020.
- Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
- High-fidelity guided image synthesis with latent diffusion models. In CVPR, 2023.
- Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
- Video inpainting by jointly learning temporal structure and spatial details. In AAAI, 2019.
- Video-to-video synthesis. NeurIPS, 2018.
- Image quality assessment: from error visibility to structural similarity. TIP, 2004.
- Towards robust referring image segmentation. arXiv preprint arXiv:2209.09554, 2022.
- Betrayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation. ICCV, 2023a.
- Towards open vocabulary learning: A survey. arXiv preprint arXiv:2306.15880, 2023b.
- Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In ICCV, 2023c.
- Smartbrush: Text and shape guided object inpainting with diffusion model. In CVPR, 2023.
- YouTube-VOS: Sequence-to-sequence video object segmentation. In ECCV, 2018.
- Joint feature learning and relation modeling for tracking: A one-stream framework. In ECCV, 2022.
- Inst-inpaint: Instructing to remove objects with diffusion models. arXiv preprint arXiv:2304.03246, 2023.
- A-ViT: Adaptive tokens for efficient vision transformer. In CVPR, 2022.
- Free-form image inpainting with gated convolution. In ICCV, 2019.
- Inpaint anything: Segment anything meets image inpainting. arXiv preprint arXiv:2304.06790, 2023.
- Contextual object detection with multimodal large language models. arXiv preprint arXiv:2305.18279, 2023.
- Learning joint spatial-temporal transformations for video inpainting. In ECCV, 2020.
- Flow-guided transformer for video inpainting. In ECCV, 2022.
- Magicbrush: A manually annotated dataset for instruction-guided image editing. In NeurIPS, 2023a.
- Adding conditional control to text-to-image diffusion models. In ICCV, 2023b.
- Places: A 10 million image database for scene recognition. TPAMI, 2017.
- ProPainter: Improving propagation and transformer for video inpainting. In ICCV, 2023.
- Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.