Task-level composition
ISpark Video
Fusion Model
One task. Specialized models. Local and global consistency.
VFM reads the task, assigns each span to the strongest generator, aligns neighboring spans locally, then checks the full sequence globally.
Model contributionEach model owns the span it does best.
Local consistencyNeighboring spans share motion and visual state.
Global consistencyThe complete sequence is checked as one task.
- 01Understand
- 02Compose
- 03Execute
- 04Evaluate
- 05Repair
- 06Deliver
Temporal completion
Fill the missing time. Keep the observed anchors.
Start with 5 seconds. VFM keeps 5 observable anchors, fills the 10 missing positions, checks continuity, and returns one 15-second sequence.
A five-second sparse input keeps five observed anchors. Ten generated slots fill four gaps, then a continuity scan validates one fifteen-second output rail.
Output refinement
Refine spatial detail and motion continuity.
Once the sequence exists, VFM can refine spatial detail and motion continuity. These studies show the same task from two measurable views.