The idea
A multi-phone recording becomes an editable scene with people, objects, lighting, audio, and time. Hide a person, move the cake, relight the room, change the virtual camera, and export a new video.
How it works
- Several phones record a moment from different viewpoints and provide evidence about surfaces hidden from any single camera.
- Reconstruct an editable scene with named people and objects, lighting and audio tracks, and a timeline; areas hidden from every camera remain unknown.
- Change scene elements or the virtual camera, then render the edit to MP4.
AI’s role: Help label people and objects and translate edit requests into scene operations; segmentation, identity tracking, geometry, and rendering must remain grounded in the recorded views.
First demonstration
Record one person walking through a room from several phones. Try Remove Person while moving the virtual camera slightly, then relight the person and inspect where the recorded views support the result.
What to solve next
Can segmentation, identity, hidden geometry, and time-coherent rendering stay stable enough for edits as the virtual camera moves?
