The idea
Phones around a court join as a temporary live sensor array, combining their camera views and motion data to reconstruct a scene that a viewer can explore from nearby virtual viewpoints.
How it works
- Several phones stream video with timestamps, ARKit camera poses, and motion data; depth and calibration observations can be included where available.
- A receiver combines the streams into a live 4D scene, with usable viewpoints limited by camera coverage, occlusion, network conditions, and compute.
- A viewer can move a virtual camera between captured viewpoints, such as from courtside to behind the basket.
AI’s role: Help identify the subject and interpret capture quality; synchronization, calibration, reconstruction, and rendering depend on measured sensor data and remain research problems.
First demonstration
Have four iPhones capture one dancer in an approximately 3×3 m space, stream to one Mac, and view the result on a fifth device. Test a supported camera move of about ±30 degrees, then measure actual end-to-end latency and image quality.
What to solve next
Can streams stay synchronized with sparse coverage, occlusion, and changing network and compute load, and how much useful viewpoint movement does the setup support?
