option
Home
Flash News
Content
CarlGarcia
CarlGarcia
October 9, 2026

Melbourne and UNSW researchers introduce VoxMem, a new benchmark exposing critical flaws in current voice AI memory evaluation. Unlike prior methods that ignore audio nuances, VoxMem isolates conversation length to test recall of speaker identity, tone, and background sounds across 34,743 audio instances. Testing 15 mainstream models reveals that none exceed 40% accuracy at 32K context length. Models significantly struggle with paralinguistic cues and environmental sounds, with performance declining as history length increases, highlighting a major gap in long-form audio understanding capabilities.

Comments (0)
0/300
OR