
The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally.
Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA’s latest open models, software and developer tools, plus the accelerated computing, libraries and educational resources that help users get started.
It’s shaping up to be a big month for agents. Follow along for the latest developments in this special-edition NVIDIA Local AI blog series, with new updates added over the coming weeks.
NVIDIA recently released Cosmos 3 Edge , a 4-billion-parameter open world model for robotics, autonomous vehicles and vision AI. At a quarter the size of Cosmos 3 Nano, it runs on device on NVIDIA DGX Spark and NVIDIA Jetson.
MiniMax-H3 is a 33-billion-parameter open weights model that generates video and natively synchronized stereo audio from text, images, video, audio or a mix of all four. Creators can access it through ComfyUI and run it locally with checkpoints optimized for NVIDIA GPUs.
Poolside AI launched Laguna S 2.1 , a 118-billion-parameter open weight agentic coding model that works through hours-long tasks. An NVFP4 checkpoint enables developers to run the model locally on a single NVIDIA DGX Spark with lower compute and memory requirements, without sacrificing accuracy.
DeepSeek recently refreshed DeepSeek-V4-Flash , a 284-billion-parameter MoE model with 13 billion active parameters and a 1 million-token context window. Developers can run it locally on an NVIDIA DGX Station using community-built GGUF versions available today.
Thinking Machines Lab’s Inkling-Small is a frontier-class, open weight multimodal model with native reasoning across text, images and audio, plus adjustable thinking effort. Trained on NVIDIA GB300 NVL72, the 276-billion-parameter model activates just 12 billion parameters per token and runs on a single DGX Station or two DGX Spark systems. An NVFP4 checkpoint optimized for NVIDIA Blackwell is available on Hugging Face .
Unsloth is launching Unsloth Desktop on Monday. The product brings local model inference, image and video diffusion, fine-tuning, agent integrations, web research and code execution into one fully open source desktop app. The launch materials position it as the first desktop app that both trains and runs AI models locally.
Unsloth Desktop brings local AI training and inference into one fully open source desktop app. It’s the first desktop app to train and run AI models locally.
Alibaba recently released Wan-Animate-2, a 14-billion-parameter open weight model that transfers motion and facial expressions from a driving video onto a static character image — human, cartoon, robot or animal. With day-zero support in ComfyUI, it generates up to 16x faster on NVIDIA RTX PRO 5000 Blackwell (48GB) and 26x faster on an NVIDIA RTX 5090 compared to Apple M3 Ultra.
LTX-2.5 is the latest state-of-the-art, open-world video generation model from LTX’s leading research team, providing the foundation for use cases across media and entertainment, robotics and real-time generation, and delivering higher-quality video, stronger prompt adherence and more consistent characters, scenes and voices across generations.
New multishot support enables creators to generate sequences spanning multiple cuts while maintaining continuity, while an upgraded diffusion video decoder improves visual fidelity and reduces artifacts. Creators can also refine existing footage with generative edits, making it easier to fix details and maintain a polished look without starting over.
The new prompt enhancer, featuring Gemma4 E2B and a custom Gemma4 12B text encoder, provides remarkable prompt adherence, giving far greater control over the output.
A new diffusion video decoder complements the variational autoencoder video decoder, improving video quality by giving a second decode path for higher-quality final renders.
LTX-2.5 is optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station systems. On an NVIDIA RTX 6000 PRO GPU, LTX-2.5 delivers up to 20% faster performance and 40% memory savings. Creators can use these NVFP4, FastVideo and ComfyUI enhancements locally through a ready-to-use ComfyUI workflow for text, image and video-to-video generation.
Learn more about LTX-2.5 to learn how to get started with high-quality, local video generation on NVIDIA hardware.
Today, Meta released Muse Glimmer, a 30-billion-parameter, dense, open weight model with a 120K+ context window. It’s purpose-built for coding and local agentic AI.
Optimized for NVIDIA GeForce RTX PCs, NVIDIA DGX Spark, DGX Station and NVIDIA Jetson , Muse Glimmer delivers over 200 tokens per second on RTX 5090 , enabling always-on agents to process data locally and work through complex, multistep tasks on a single system.
Its dense architecture and hybrid attention help keep processing and memory demands manageable as agents take on longer tasks, use tools and maintain context across multiple steps.
Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a PC with a single consumer GPU, enabling use cases that range from local agents and function calling to local coding and LLM-as-a-judge evaluation.
AI enthusiasts and developers can build agents using NemoClaw and fine-tune Muse Glimmer locally with NVIDIA NeMo Automodel , using private or specialized data while keeping it on the system. Muse Glimmer is designed for:
Custom agents: Adapt the model for specific tasks, tools, data and areas of expertise.
Private data processing: Read, summarize and act on local files, documents, emails and messages.
Credential handling: Use application programming interface keys, authentication tokens and other sensitive information while keeping inference on the device.
Multistep tasks: Complete many sequential tool calls and recover from errors or unexpected results.
Long-running workflows: Break larger projects into steps, track progress and resume interrupted sessions with context intact.
Developers can run Muse Glimmer with popular inference frameworks including vLLM for text generation, image, reasoning and tool use, or llama.cpp for text and image workloads using BF16 and quantized GGUF checkpoints. llama.cpp also supports DFlash speculative decoding to accelerate generation.
Running on a single NVIDIA RTX 5090 , Muse Glimmer enables always-on agents to work across private files, applications and communications while reducing reliance on cloud-based inference.
Learn more about Meta’s Muse Glimmer and check out the NVIDIA tech blog to learn how to get started with local AI and fine-tuning on NVIDIA hardware.
New open models such as GLM 5.2 and DeepSeek V4 Flash are bringing cloud-level intelligence to local systems, enabling advanced workloads like coding and research agents. Running these larger models, however, can require multiple GPUs or DGX Spark systems working together.
NVIDIA Sync app updates make it easy to cluster multiple DGX Spark systems together, providing the memory capacity, inference performance and training throughput needed to run larger models, faster.
Available for Windows and macOS, NVIDIA Sync automatically detects connected systems, provides private and secure remote access through Tailscale and lets developers launch applications across one or more DGX Spark systems without manual networking or copied terminal commands.
The Cluster Assistant in NVIDIA Sync, automates configuring two or more DGX Spark systems as a high-speed cluster. Developers connect the systems through their NVIDIA ConnectX-7 ports, and NVIDIA Sync configures the network, routes workloads across nodes and monitors system health.
To get started with agentic AI on DGX Spark, check out playbooks on NemoClaw , OpenClaw , Hermes Agent and OpenShell .
See notice regarding software product information.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/#primary
- https://blogs.nvidia.com/blog/author/nvidiawriters/
- https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/#disqus_thread
- NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents
- Chinese farmer kills 25 acres of crops after following AI-generated weed and pest control advice — farmer trusted pesticide recipe after months of successful ad
- Noctua finds more than half of tested PC cases misstate CPU cooler clearances — hands-on checks reveal errors ranging from -3.5mm to +10mm, internal compatibili
- US lawmaker wants gov't to enforce regulation to ensure 'chipmakers conduct adequate due diligence on their customers' — House member calls for Biden-era export
- Windows 11's built-in weather app hogs more than 1.2 gigabytes of RAM just to tell the forecast — memory-sucking web wrapper filled with ads masquerades as an a
Informational only. No financial advice. Do your own research.