UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced
UniWorld-View, open-sourced by Peking University + Rabbitpre + Pengcheng Laboratory, generates novel-view videos along any camera trajectory from a single image/video — it tops Fei-Fei Li's team's WorldScore leaderboard, with code and weights fully open under Apache-2.0 and one-click downloads.


UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced
UniWorld-View, open-sourced by Peking University + Rabbitpre + Pengcheng Laboratory, generates novel-view videos along any camera trajectory from a single image/video — it tops Fei-Fei Li's team's WorldScore leaderboard, with code and weights fully open under Apache-2.0 and one-click downloads.
If you work in 3D reconstruction, film VFX, spatial computing, or game previsualization, "generating arbitrary-viewpoint video from a single image" has probably been tormenting you for a while. UniWorld-View is a world model open-sourced jointly by Peking University (Prof. Li Yuan's group), Rabbitpre, and Pengcheng Laboratory: feed it a casually shot monocular video or even a single image, and it outputs novel-view video (Novel View Synthesis, NVS) along any camera trajectory + 4D reconstruction. It ranks first on the WorldScore leaderboard maintained by Fei-Fei Li's team, its code and weights are fully open under Apache-2.0, and it adapts to domestic Ascend compute. This piece walks you through running it locally.
What UniWorld-View Is
One-line positioning: a single image/video → novel-view video generator for any camera trajectory.
It tackles the core difficulty of NVS (Novel View Synthesis) — you only see the scene from one viewpoint, and the model has to imagine the video seen from any other angle, geometrically and visually plausible and consistent. Unlike image-to-video models, UniWorld-View emphasizes geometric consistency and can directly serve 4D reconstruction, spatial asset production, and virtual cinematography previz.
The official words: "ranked 1st on the WorldScore Leaderboard (by Stanford Prof. Fei-Fei Li's Team)" — that is, first place on the world-model evaluation leaderboard maintained by Fei-Fei Li's team.
Technical Highlights
Two key designs determine its generation quality:
- Occlusion-aware Point Cloud Rendering: the input image/video is first inverted into a point cloud carrying occlusion information, and new viewpoints are rendered from that point cloud, avoiding the "free hallucination" that causes geometry to clip through itself.
- Dual-stream Conditioning: conditions split into a geometry stream + an appearance stream — the geometry stream runs through the Ref-DiT module to control structure, the appearance stream handles color and texture, and both streams feed the generation backbone simultaneously, avoiding "shape right, texture flying off" or the reverse.
This geometry/appearance decoupling is the structural reason it can take first place on WorldScore.
Before You Start
- GPU: the official recommendation is ≥ 60GB of VRAM for smooth inference (an A100 80G / H100 is the ideal setup). Smaller cards can try, but it's not the recommended path.
- Python environment: configure per the repo README (PyTorch + matching CUDA).
- Storage: the weights are large; reserve enough disk space.
- Domestic Ascend compute: the repo states it supports Ascend adaptation; see the README for details.
Project repository: https://github.com/PKU-YuanGroup/UniWorld-View (Apache-2.0) Model weights: https://huggingface.co/Drexubery/UniView
Step 1: Clone the Repository
git clone https://github.com/PKU-YuanGroup/UniWorld-View.git
cd UniWorld-ViewStep 2: One-Click Weight Download
The repo ships a script that pulls all checkpoints from HuggingFace:
bash checkpoints/download_hf.shOnce the download finishes, the weights land in the checkpoints/ directory and load automatically at runtime.
Step 3: Run Inference
Command-line inference (good for batch jobs):
bash run_infer.shGradio visual interface (good for interactive tuning):
bash run_app.shOnce it's running, upload an image or a monocular video, specify the camera trajectory you want (push in, pull out, orbit, lateral tracking, etc.), and wait for generation to finish to see the novel-view video.
💡 Tip: When VRAM is tight, turn the generation resolution and frame count down in the Gradio interface for iterative validation first, confirm the camera trajectory is sensible, then move up to high-quality settings for the final run. Geometric correctness deserves verification before texture detail.
Verification Results
The output is a video showing the input scene from the entirely new viewpoint defined by your specified camera trajectory. Three criteria for success:
- Geometric consistency: objects' relative positions and shapes show no clipping or drift.
- Appearance consistency: colors, textures, and lighting stay consistent with the input image.
- Controllable trajectory: the camera follows your specified trajectory strictly, with no accidental whip-pans.
Use Cases
- Spatial computing / VR / AR assets: single-viewpoint footage → multi-viewpoint navigable assets, skipping a full 360° shoot.
- Film VFX and virtual cinematography previz: from one live-action shot, quickly generate "what another camera would see" for pre-editing and composition checks.
- 4D reconstruction / digital twins: feed NVS results into a 4D reconstruction pipeline to get time-varying geometry.
- Game previsualization and level exploration: from one concept image, quickly generate multi-viewpoint exploration videos to validate level design.
FAQ
- Out of VRAM (OOM): the official recommendation is ≥ 60GB. For 40GB-class cards, lower resolution and frame count; consumer cards (24GB) are not the recommended target hardware for this project right now.
- Download script slow / timing out:
checkpoints/download_hf.shpulls from HuggingFace by default; on networks in China, use a mirror or proxy. Weights address: https://huggingface.co/Drexubery/UniView - Geometry drifts in the output: check whether the input image is too planar or texture-poor — the geometry stream needs enough texture features to align on, and plain solid-color walls are the enemy.
- Ascend adaptation: the repo states Ascend support; see the relevant README section for exact dependency versions and build options.
Final Thoughts
UniWorld-View's engineering completeness lives up to the words "fully open-sourced": HuggingFace weights + a one-click download script + dual command-line / Gradio entry points + the Apache-2.0 license — you can start today. If your work touches 3D / film / spatial computing, pull it down and run run_app.sh once, draw a camera trajectory yourself, and see the result — far more intuitive than reading the paper.
Official resources:
- GitHub: https://github.com/PKU-YuanGroup/UniWorld-View
- HuggingFace weights: https://huggingface.co/Drexubery/UniView
- Paper: see the link at the top of the GitHub README