MonoFusion reconstructs dynamic 4D scenes from sparse-view videos captured by four static cameras. It aligns monocular geometry and motion estimates across views and time, using depth priors and semantic features to initialize reconstruction. The paper evaluates novel-view synthesis on Panoptic Studio and ExoRecon, a subset of Ego-Exo4D. The input is synchronized multi-camera video with known camera parameters; monocular fusion describes the reconstruction strategy.