Engineers have created a generative world model called Cosmos-H-Dreams that simulates surgical robotics in real time. Artificial intelligence video generators usually produce passive clips after long processing delays rather than reacting immediately to human steering. The new architecture converts motion signals from a robot controller into matching video frames of bending tissue with minimal delay.
Standard physics engines struggle to render soft human organs accurately because wet, yielding tissue deforms in unpredictable visual patterns. The researchers trained a teacher video model on surgical recordings and distilled its knowledge into a compact student model using a technique called Self Forcing. This student model functions like a flight simulator for surgeons, rendering each progressive step only after receiving the previous movement command. By computing frames causally forward in sequence, the software generates continuous visual feedback.
The development team tested Cosmos-H-Dreams by connecting the simulation engine to multiple input interfaces. Powered by a single Nvidia RTX PRO 6000 Blackwell graphics card, the distilled model streamed video at about one hundred sixty frames per second. Test operators successfully directed simulated tools using web browsers, Meta Quest virtual headsets, and commercial Versius surgical consoles.
The team released the software as an open system for surgical education, synthetic data generation, and intraoperative decision support. Human operators and automated control policies can now manipulate virtual instruments and see tissue reactions without animal or cadaver trials.
