FloMo: 3D Scene Flow as a World-Action Model Intermediate for Learning from Human Video
Jeremy A. Collins*, Namra Patel*, Ayush Agarwal*, Shitij Govil*, Animesh Garg
In Submission
Paper |
Website
Given an image and text prompt, FloMo jointly predicts dense 3D motion (scene flow) and robot actions using a pretrained 5B-parameter video generation model using flow-matching. To tokenize scene flow, we transform instantaneous 3D velocities into RGB videos that are encoded by the pretrained video backbone's video encoder.