H A R M O N I
Aligning Human and Scene Priors for
Multi-View 4D Reconstruction

1 Seoul National University     2 NAVER Cloud

TL;DR: Given monocular or multi-view video, HARMONI reconstructs cameras, scene, and humans with identities in a feed-forward pass by aligning complementary human and scene priors.

HARMONI teaser: reconstructed humans and scene from multi-view and monocular video, and a comparison with UniSH and Human3R.

Abstract

Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independently trained human and scene priors often produces misalignment in scale and depth. We observe that the two priors have complementary strengths. The scene prior provides consistent depth but approximate scale, while the human prior provides fixed body scale but less reliable depth. Motivated by this observation, we present HARMONI, a feed-forward framework that reconstructs cameras, scene, and humans with identities from monocular or multi-view video without test-time optimization. We introduce bidirectional anchoring, in which scene depth guides human placement, while human keypoints calibrate the scene scale to match the human body. While most existing approaches target monocular inputs and multi-view methods rely on optimization or re-identification, our approach naturally extends to multiple views. Building on bidirectional anchoring, we introduce a multi-view fusion module that merges per-view estimates, while reducing the influence of unreliable views and visually similar individuals. Experiments show that our framework outperforms previous human-scene methods in global motion and multi-view pose estimation by up to 28% and 64%, while running 28× faster than optimization-based approaches.

Interactive Viewer

Explore the reconstructed humans and scene interactively. (Points and frames are downsampled.)

0 / 9

Multi-View Results

Given multi-view videos, HARMONI reconstructs coherent human meshes and the surrounding scene in diverse settings.

Monocular Results

HARMONI can also be applied to a monocular setup.

BibTeX


    @misc{kim2026harmoni,
          title={HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction}, 
          author={Sangmin Kim and Minhyuk Hwang and Geonho Cha and Dongyoon Wee and Jaesik Park},
          year={2026},
          eprint={2603.12789},
          archivePrefix={arXiv},
          primaryClass={cs.CV},
          url={https://arxiv.org/abs/2603.12789}
    }