NVIDIA Unveils Nemotron-Labs-Audex-30B-A3B Unified Audio Intelligence Model

As multimodal large models evolve rapidly, audio processing capabilities are frequently compromised—many models improve audio understanding at the expense of text logic. To address this, NVIDIA researchers have introduced Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text large language model designed to bridge this gap.
Audex employs a streamlined and efficient design, built on a robust pure-text Mixture of Experts (MoE) architecture. By utilizing a single Transformer decoder, it processes text and quantized audio tokens simultaneously. This approach projects audio inputs into the text embedding space, ensuring seamless integration with existing LLM infrastructure for multimodal tasks and achieving true deep integration.
Training involved extensive datasets comprising 157.4 billion audio tokens and 320.5 billion text tokens. Through multi-stage supervised training, pure-text Cascade Reinforcement Learning, and multi-domain in-policy knowledge distillation, Audex achieves industry-leading performance in audio understanding, speech recognition, translation, and generation. Crucially, it preserves the original LLM’s core capabilities in reasoning, alignment, knowledge retrieval, and long-text processing with minimal degradation.
As an open-source model, Audex marks a significant milestone for the voice technology sector. Moving beyond theoretical research, it offers a mature solution for developers to evaluate and deploy directly. For those managing complex audio interactions, Audex provides a balanced option that enhances performance while opening new avenues for multimodal intelligence research.
Related article
Claude Opus 5.2 Night Gray Launches with Faster Response, Solving Laziness Issue
Opus 5.2 quietly launched this morning, prompting many developers to notice that the updated model, Claude Opus 5.2, has begun a limited rollout within Claude Code.Yesterday evening, X users observed that invoking Opus 5 in Claude Code yielded perfor
Fireworks AI Unveils FireRouter with Opus: Cuts Encoding Costs by 57% with Minimal Accuracy Drop
Fireworks AI has introduced FireRouter with Opus, the industry’s first cache-aware routing system tailored for the Claude Opus series. Now accessible via a serverless endpoint, this independent routing model underwent over a month of internal A/B tes
How to fix core web vitals for mobile seo
Boost Local SEO: Build Your Google Entity Cloud Drive StacksTable of Contents:IntroductionUnderstanding Google Entity Cloud StackingKey Benefits of Google Entity Cloud StackingGetting Started with Google Entity Cloud Stacks 4.1. Option 1: Acquire Age
Related Special Topic Recommendations
Comments (0)
0/500

As multimodal large models evolve rapidly, audio processing capabilities are frequently compromised—many models improve audio understanding at the expense of text logic. To address this, NVIDIA researchers have introduced Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text large language model designed to bridge this gap.
Audex employs a streamlined and efficient design, built on a robust pure-text Mixture of Experts (MoE) architecture. By utilizing a single Transformer decoder, it processes text and quantized audio tokens simultaneously. This approach projects audio inputs into the text embedding space, ensuring seamless integration with existing LLM infrastructure for multimodal tasks and achieving true deep integration.
Training involved extensive datasets comprising 157.4 billion audio tokens and 320.5 billion text tokens. Through multi-stage supervised training, pure-text Cascade Reinforcement Learning, and multi-domain in-policy knowledge distillation, Audex achieves industry-leading performance in audio understanding, speech recognition, translation, and generation. Crucially, it preserves the original LLM’s core capabilities in reasoning, alignment, knowledge retrieval, and long-text processing with minimal degradation.
As an open-source model, Audex marks a significant milestone for the voice technology sector. Moving beyond theoretical research, it offers a mature solution for developers to evaluate and deploy directly. For those managing complex audio interactions, Audex provides a balanced option that enhances performance while opening new avenues for multimodal intelligence research.
Claude Opus 5.2 Night Gray Launches with Faster Response, Solving Laziness Issue
Opus 5.2 quietly launched this morning, prompting many developers to notice that the updated model, Claude Opus 5.2, has begun a limited rollout within Claude Code.Yesterday evening, X users observed that invoking Opus 5 in Claude Code yielded perfor
Fireworks AI Unveils FireRouter with Opus: Cuts Encoding Costs by 57% with Minimal Accuracy Drop
Fireworks AI has introduced FireRouter with Opus, the industry’s first cache-aware routing system tailored for the Claude Opus series. Now accessible via a serverless endpoint, this independent routing model underwent over a month of internal A/B tes
How to fix core web vitals for mobile seo
Boost Local SEO: Build Your Google Entity Cloud Drive StacksTable of Contents:IntroductionUnderstanding Google Entity Cloud StackingKey Benefits of Google Entity Cloud StackingGetting Started with Google Entity Cloud Stacks 4.1. Option 1: Acquire Age





Home






