Will it run?
Models

Nvidia researchers unveil Physis-Lang, a physics-grounded language layer that pushes Cosmos 3 past Veo 3.1

By Rae Whitlock Clawpit staff
Nvidia researchers unveil Physis-Lang, a physics-grounded language layer that pushes Cosmos 3 past Veo 3.1

Video world models can produce convincing clips yet still violate basic physics — butter spreads like paint, balls pass through walls. A team from Nvidia, MIT and the University of Oxford argues the fix will not come from more visual, latent or numerical signals, but from language itself. Their framework, Physis-Lang, treats physical language as a shared, optimisable representation: the same text drives data collection, training and inference. On the Physics-IQ Verified leaderboard dated 29 September 2026, Physis-Lang atop Cosmos3-Super takes first place with 48.2 ± 1.4; the Nano variant ranks second at 43.3 ± 1.5.

Standard captions describe what happens, not why. "Butter melts when temperature rises" says nothing about heat transfer or gravity. Physis-Lang adds a physics_reasoning field to every base caption — entities, causes, interactions, governing principles, temporal evolution and outcomes. The pipeline also generates a scene-specific physics_negative_prompt describing implausible outcomes (a stone floating on water) that serves as a negative condition at inference time. The evolving-caption loop keeps the writer (GPT-5.5) frozen and updates only its prompt. Gemini-3.1-Pro acts as a physics-aware critic, scoring on two dimensions: precision (every atomic claim checked against the video) and recall (every human-verified claim must appear explicitly or be derivable from the caption). An evolution agent reads the scores and claim-level failures and rewrites the prompt. Validation runs on PhysCapBench, a new benchmark of 246 videos and 3,794 verified claims. Caption F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9; the path was not monotonic — iteration 2 made captions over-cautious and dropped F1 to 76.28. By iteration 9 every visible causal step was required, and frame sampling increased from 2 to 4 fps.

A GPT-5.5-based diagnostic agent maps generated-video failures to physical categories — rigid-body motion, collision, fluid dynamics. That deficiency profile is cross-referenced against physics tags in a large video gallery so retrieval targets physical content rather than visual appearance. The final training set contains 183 videos: 71 filtered from WISA-80K plus 112 retrieved. Retrieval alone added 3.01 points on average across three benchmarks; on VideoPhy-2, chemical and thermal processes each gained 8.00 points.

Fine-tuning uses LoRA on the attention projections with no architecture or objective change. Against Google's Veo 3.1, Cosmos3-Nano with Physis-Lang scores: PhyGenBench 71.04 vs 65.63; Physics-IQ Verified 43.41 vs 34.99; PhyGround 69.90 vs 69.24; VideoPhy-2 full set 68.02 vs 68.87, hard split 62.36 vs 58.43. Gains replicate across backbones: +7.05 on Wan2.1-14B, +3.24 on Cosmos3-Edge-4B, +6.22 on Cosmos3-Nano-16B and +5.02 on Cosmos3-Super-64B. General quality held steady on VBench-I2V, where Cosmos3-Nano moved from 88.32 to 88.69. Prompting alone helps: physics reasoning with negative prompts lifted a frozen Cosmos3-Nano on PhyGenBench from 61.67 to 67.29.

The team distilled the commercial pipeline into two Qwen3-VL-4B-Instruct models: PhysThinker-C for captioning and PhysThinker-U for prompt enrichment. On Wan2.1-14B the commercial pipeline delivered +7.05 at roughly $24,120 in API costs. Swapping in PhysThinker-C kept +6.76 at about $120. A fully local setup cost zero and still added +4.76.