Papers
Topics
Authors
Recent
Search
2000 character limit reached

Towards Language-Driven Video Inpainting via Multimodal Large Language Models

Published 18 Jan 2024 in cs.CV | (2401.10226v2)

Abstract: We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that depend on manually labeled binary masks, a process often tedious and labor-intensive. We present the Remove Objects from Videos by Instructions (ROVI) dataset, containing 5,650 videos and 9,091 inpainting results, to support training and evaluation for this task. We also propose a novel diffusion-based language-driven video inpainting framework, the first end-to-end baseline for this task, integrating Multimodal LLMs to understand and execute complex language-based inpainting requests effectively. Our comprehensive results showcase the dataset's versatility and the model's effectiveness in various language-instructed inpainting scenarios. We will make datasets, code, and models publicly available.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (67)
  1. Instructpix2pix: Learning to follow image editing instructions. In CVPR, 2023.
  2. Language models are few-shot learners. NeurIPS, 2020.
  3. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  4. Free-form video inpainting with 3d gated convolution and temporal patchgan. In ICCV, 2019.
  5. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  6. Video inpainting with short-term windows: application to object removal and error concealment. TIP, 2015.
  7. Actor and action video segmentation from a sentence. In CVPR, 2018.
  8. Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041, 2023.
  9. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022.
  10. Denoising diffusion probabilistic models. NeurIPS, 2020.
  11. Proposal-based video completion. In ECCV, 2020.
  12. Gqa: A new dataset for real-world visual reasoning and compositional question answering. CVPR, 2019.
  13. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023.
  14. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  15. Learning blind video temporal consistency. In ECCV, 2018.
  16. Copy-and-paste networks for deep video inpainting. In ICCV, 2019.
  17. Short-term and long-term context aggregation network for video inpainting. In ECCV, 2020a.
  18. Mimic-it: Multi-modal in-context instruction tuning. arXiv preprint arXiv:2306.05425, 2023a.
  19. Recurrent feature reasoning for image inpainting. In CVPR, 2020b.
  20. MAT: Mask-aware transformer for large hole image inpainting. In CVPR, 2022a.
  21. Transformer-based visual segmentation: A survey. arXiv pre-print, 2023b.
  22. Towards an end-to-end framework for flow-guided video inpainting. In CVPR, 2022b.
  23. Image inpainting for irregular holes using partial convolutions. In ECCV, 2018.
  24. Partial convolution for padding, inpainting, and image synthesis. TPAMI, 2022.
  25. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023a.
  26. Fuseformer: Fusing fine-grained information in transformers for video inpainting. In ICCV, 2021a.
  27. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023b.
  28. Deep learning face attributes in the wild. In ICCV, 2015.
  29. Swin Transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021b.
  30. Repaint: Inpainting using denoising diffusion probabilistic models. In CVPR, 2022.
  31. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021.
  32. AdaViT: Adaptive vision transformers for efficient image recognition. In CVPR, 2022.
  33. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022.
  34. OpenAI. Gpt-4 technical report, 2023.
  35. Context encoders: Feature learning by inpainting. In CVPR, 2016.
  36. A benchmark dataset and evaluation methodology for video object segmentation. In CVPR, 2016.
  37. Detgpt: Detect what you need via reasoning. arXiv preprint arXiv:2305.14167, 2023.
  38. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  39. DynamicViT: Efficient vision transformers with dynamic token sparsification. In NeurIPS, 2021.
  40. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  41. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 2022.
  42. Urvos: Unified referring video object segmentation network with a large-scale benchmark. In ECCV, 2020.
  43. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  44. High-fidelity guided image synthesis with latent diffusion models. In CVPR, 2023.
  45. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  46. Video inpainting by jointly learning temporal structure and spatial details. In AAAI, 2019.
  47. Video-to-video synthesis. NeurIPS, 2018.
  48. Image quality assessment: from error visibility to structural similarity. TIP, 2004.
  49. Towards robust referring image segmentation. arXiv preprint arXiv:2209.09554, 2022.
  50. Betrayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation. ICCV, 2023a.
  51. Towards open vocabulary learning: A survey. arXiv preprint arXiv:2306.15880, 2023b.
  52. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In ICCV, 2023c.
  53. Smartbrush: Text and shape guided object inpainting with diffusion model. In CVPR, 2023.
  54. YouTube-VOS: Sequence-to-sequence video object segmentation. In ECCV, 2018.
  55. Joint feature learning and relation modeling for tracking: A one-stream framework. In ECCV, 2022.
  56. Inst-inpaint: Instructing to remove objects with diffusion models. arXiv preprint arXiv:2304.03246, 2023.
  57. A-ViT: Adaptive tokens for efficient vision transformer. In CVPR, 2022.
  58. Free-form image inpainting with gated convolution. In ICCV, 2019.
  59. Inpaint anything: Segment anything meets image inpainting. arXiv preprint arXiv:2304.06790, 2023.
  60. Contextual object detection with multimodal large language models. arXiv preprint arXiv:2305.18279, 2023.
  61. Learning joint spatial-temporal transformations for video inpainting. In ECCV, 2020.
  62. Flow-guided transformer for video inpainting. In ECCV, 2022.
  63. Magicbrush: A manually annotated dataset for instruction-guided image editing. In NeurIPS, 2023a.
  64. Adding conditional control to text-to-image diffusion models. In ICCV, 2023b.
  65. Places: A 10 million image database for scene recognition. TPAMI, 2017.
  66. ProPainter: Improving propagation and transformer for video inpainting. In ICCV, 2023.
  67. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
Citations (11)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.