What Netflix released
On April 3, 2026, Netflix published VOID (Video Object and Interaction Deletion) on Hugging Face and GitHub — the company's first publicly available open-source AI model. VOID is a video inpainting model, meaning it removes specified objects from video footage and fills in the resulting gaps. But it differs from existing tools in one significant way: it accounts for physical interactions.
Standard video inpainting tools remove an object and fill the space with plausible background pixels. VOID goes further. When you remove a person who is holding a guitar, for instance, VOID doesn't just erase the person — it also simulates what would happen to the guitar without the person holding it. The guitar falls. Objects on a table that were being held in place slide or topple. The model attempts to reconstruct physically plausible outcomes for everything the removed object was interacting with.
Netflix released the model weights under an Apache 2.0 license, making it available for both research and commercial use. The code, weights, and a demonstration interface are all publicly accessible.
How VOID works: quadmask conditioning
VOID is built on top of CogVideoX-Fun-V1.5-5b-InP, a video generation model, and fine-tuned specifically for interaction-aware video inpainting. The key technical contribution is what the Netflix team calls "quadmask conditioning" — a four-value mask system that encodes different semantic regions of the scene.
The four mask values represent:
- Primary object — the item being removed
- Overlap regions — areas where the removed object intersects with other objects
- Affected regions — objects that would be physically affected by the removal (items that would fall, slide, or shift)
- Background — areas to preserve unchanged
This four-way decomposition gives the model explicit information about physical dependencies in the scene, rather than requiring it to infer all relationships from pixel data alone. The mask generation pipeline uses SAM2 (Segment Anything Model 2) combined with Gemini to automatically create these quadmasks, though Netflix also provides a GUI editor for manual refinement.
Inference runs in two passes. Pass 1 handles the core inpainting and is sufficient for most clips. Pass 2 adds optical flow-warped latent initialization — essentially using motion information from the original video to improve temporal consistency across longer sequences.
Training approach
The model was trained on paired counterfactual videos from two synthetic sources: HUMOTO, which contains human-object interactions rendered in Blender with physics simulation, and Kubric, which uses Google Scanned Objects for object-only interaction scenarios. Training ran on 8x A100 80GB GPUs using DeepSpeed ZeRO Stage 2. By training on synthetic data where the "ground truth" outcome (what happens when an object is removed) is known, the team avoided the need for manually annotated real-world video pairs.
Benchmarks and comparisons
Netflix reports results from a human preference study involving 25 evaluators across multiple scenarios. According to the study, VOID was preferred 64.8% of the time, with Runway's inpainting coming in second at 18.4%. The remaining preferences were split across other tools.
These numbers come with caveats. The evaluation pool of 25 people is small, and the scenarios were selected by the researchers. The model's strength is specifically in interaction-aware removal — scenes where physical consequences matter. For simple object removal in static scenes without physical dependencies, the performance gap over existing tools is likely smaller.
The paper, published as a preprint on arXiv (2604.02296), provides additional quantitative metrics, but the human preference study is the most cited result so far.
Practical implications for video editors
For professional video editors and VFX artists, VOID addresses a specific pain point: removing objects or people from footage while maintaining physical plausibility. Current commercial tools handle simple removals reasonably well but tend to produce artifacts when the removed object is physically interacting with its environment — holding things, casting dynamic shadows, pushing objects, or supporting other items.
The immediate practical applications include:
- Post-production cleanup — removing unwanted people or objects from shots without breaking the physical logic of the scene
- Privacy and compliance — erasing individuals from footage while keeping the rest of the scene coherent
- Creative experimentation — testing "what if this weren't here" scenarios during the edit process
However, the hardware requirements are steep. VOID requires a GPU with at least 40GB of VRAM (an A100 or equivalent), which puts it out of reach for most individual creators. This is a research-grade tool at the moment, not a plug-and-play solution for consumer editing software.
The Apache 2.0 license means commercial tool makers could integrate VOID's approach into their products. Whether companies like Adobe, Runway, or others adopt the quadmask technique or the model itself remains to be seen, but the open release lowers the barrier significantly.
Limitations and requirements
VOID has meaningful limitations worth noting. The model was trained primarily on synthetic data, which means its understanding of real-world physics is learned indirectly through simulated environments. Complex, multi-object interaction chains — where removing one object triggers a cascade of physical events — may not be handled reliably.
The automatic mask generation pipeline, while functional, sometimes requires manual correction. Netflix includes a GUI editor for this purpose, but the need for refinement adds friction to the workflow. The two-pass inference system also means processing is slower than single-pass alternatives.
System requirements present the largest barrier to adoption:
- GPU — 40GB+ VRAM minimum (A100 or equivalent)
- Framework — PyTorch with CUDA support
- Dependencies — SAM2 and Gemini API access for automatic mask generation
The model also does not handle audio. If a removed person was speaking or creating sound, VOID offers no mechanism for adjusting the audio track to match the visual changes.
Still, as Netflix's first public model release, VOID signals the company's willingness to contribute to the open-source AI video ecosystem. For a company that has historically kept its ML work internal, this is a notable shift — and one that gives researchers and tool builders a new foundation for physics-aware video editing.
Stop scrubbing. Start creating.
Wideframe gives your team an AI agent that searches, organizes, and assembles Premiere Pro sequences from your footage. 7-day free trial.
Frequently asked questions
VOID (Video Object and Interaction Deletion) removes objects from video footage and simulates what would physically happen in the scene without them. Unlike standard inpainting tools that just fill in backgrounds, VOID models physical consequences — if you remove a person holding a plate, the plate falls.
Yes. VOID is released under the Apache 2.0 license, which permits both research and commercial use at no cost. The model weights are available on Hugging Face and the code is on GitHub. However, running it requires a GPU with at least 40GB of VRAM.
VOID requires a GPU with 40GB or more of VRAM, such as an NVIDIA A100. It also needs PyTorch with CUDA support. For automatic mask generation, you need access to the SAM2 model and the Gemini API. This puts it firmly in professional or research territory rather than consumer hardware.
In Netflix's human preference study, VOID was preferred 64.8% of the time compared to 18.4% for Runway, specifically for interaction-aware object removal. Runway remains more accessible as a cloud-based commercial tool, while VOID requires self-hosting on high-end GPU hardware. For simple removals without physical dependencies, the gap is likely smaller.