NVIDIA centralises open models for local execution, marking August as a month of agents

A batch of announcements of seven new models and a single desktop tool place August as the month where the open ecosystem moves from talk of “local execution” to production deployment. NVIDIA leverages the Local AI blog series to push the whole stack: from world models to robotics, through video with synchronized audio, to code agents that hold a context of a million tokens, all with checkpoints aimed at Blackwell and ready to run on DGX Spark 2.1118, DGX Station or RTX 5090 GPU cards.
Robotics and vision: Cosmos 3 Edge scaled down to 4 billion parameters
Cosmos 3 Edge is an open world model at a quarter of Cosmos 3 Nano, developed for robotics, autonomous vehicles and computer vision. Practical advantage: it runs on-device on Jetson and DGX Spark, letting robotics developers test simulations and close loops without cloud. For teams building edge autonomy, the difference is between iteration cycles measured in hours versus minutes.
Video with audio: MiniMax-H3, LTX 2.5 and Wan-Animate-2 cover all angles
MiniMax-H3 brings 33 billion parameters of open weights that can generate video and synchronized stereo audio from text, image, video or audio, or any combination of the four. The approach via ComfyUI and the NVIDIA-tuned checkpoints make it accessible to creators without writing infrastructure code. LTX 2.5 adds multishot support, a sequence of cuts that preserves character, scene and voice consistency, and an improved diffusion decoder that reduces artifacts. From China, Alibaba’s Wan-Animate-2 (14 billion parameters) transfers motion and expressions from driving video to a static character, person, cartoon figure, robot or animal, with day-zero support in ComfyUI. NVIDIA cites speedups of 16× on RTX PRO 5000 Blackwell (48 GB) and 26× on RTX 5090 versus Apple M3 Ultra.
Code agents: Poolside, DeepSeek and Thinking Machines Lab push long context and sparse activation
Poolside launches Laguna S, a code-agent model that holds tasks lasting hours. NVFP4 checkpoint enables execution on a single DGX Spark without sacrificing accuracy. DeepSeek refreshes V4-Flash: a MoE model of 284 billion parameters with only 13 billion active at a time and a context window of one million tokens, running locally on a DGX Station via GGUF released by the community. Thinking Machines Lab presents Inkling-Small: 276 billion multimodal parameters with native reasoning and adaptive thinking, activating only 12 billion parameters per token and running on a single DGX Station or a pair of DGX Spark. The NVFP4 checkpoint for Blackwell is already on Hugging Face.
Unsloth Desktop: first desktop app that trains and runs in a single open source tool
Unsloth this week releases Unsloth Desktop, an open-source desktop application that unifies local inference, image and video diffusion, fine-tuning, agent integrations, web search and code execution. The company claims it is the first desktop app that trains and runs models locally without needing an external server. For researchers and developers accustomed to stitching pipelines from multiple tools, this saves significant setup time and environment maintenance.
Bottom line: infrastructure matured, only building remains
The current wave shows the bottleneck has shifted from model ordering and tailored checkpoints to teams’ ability to embed them in real products. With DGX Spark at a relatively accessible price, RTX 5090 in workstations, and day-zero coverage in ComfyUI and Hugging Face, the excuse “it doesn’t run on my machine” no longer holds water. For Israeli startups building on edge or data privacy, this month provides a ready-to-use (immediate) toolbox.