Command Palette
Search for a command to run...
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Yang Xiao Vidhyasaharan Sethu Eun-Jung Holden Ting Dang
Abstract
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
One-sentence Summary
Researchers from the University of Melbourne and the University of New South Wales introduce VoxMem, a benchmark of 3,196 evaluation instances over 34,743 spoken sessions that applies a taxonomy jointly characterizing acoustic evidence and memory operations to evaluate 15 large audio language models across 8K to 64K token contexts, finding that none exceeds 40% at 32K and that models retain lexical content far better than speaker identity, paralinguistic cues, and environmental sound, exposing gaps beyond lexical memory.
Key Contributions
- Introduces a two-dimensional taxonomy for spoken conversational memory that jointly characterizes acoustic evidence to be retained, including speech semantics, speaker identity, paralinguistic cues, and environmental sound, alongside memory operations such as information extraction, multi-session reasoning, temporal tracking, and answer refusal.
- Presents VoxMem, a multi-session benchmark with 3,196 quality-controlled instances over 34,743 spoken sessions totaling 177 hours, stratified across context budgets from 8K to 64K tokens and built from evidence, competing, and unrelated sessions across 20 topic families.
- Reports a systematic evaluation of 15 LALMs showing that no model exceeds 40% at 32K context, that lexical content is retained much better than speaker identity, paralinguistic cues, or environmental sounds, and that failure modes differ qualitatively by evidence type, such as misattribution for speaker identity and lost acoustic cues for paralinguistic information.
Introduction
Large audio language models (LALMs) are enabling longer spoken interactions across separate sessions, making long-term memory a critical requirement. However, existing spoken-memory benchmarks lack a principled taxonomy of what acoustic evidence must be retained and which memory operations are needed, and they mostly evaluate isolated recordings or continuous dialogues rather than multi-session histories. They also tend to confound history length with question difficulty and evidence type. The authors introduce VoxMem, a multi-session benchmark organized along two axes: acoustic evidence, encompassing speech semantics, speaker identity, paralinguistic cues, and environmental sound, and memory operation, comprising information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal. VoxMem contains 3,196 quality-controlled instances over 34,743 spoken sessions, spanning four context budgets from 8K to 64K tokens, and supports controlled evaluation of how spoken memory degrades as context length scales.
Dataset
VoxMem Benchmark Dataset
-
Scope and taxonomy. VoxMem is a spoken conversational memory benchmark organized around two dimensions: acoustic evidence and memory operations. Acoustic evidence covers speech semantics, speaker identity, paralinguistic cues, and environmental sound. Memory operations cover information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal. The benchmark includes 15 evaluation scenarios, leaving out information extraction over speech semantics.
-
Data composition. Each multi-session history contains three session types. Evidence sessions hold the information required to answer a question. Haystack sessions add plausible but misleading content on the same topic, preventing the model from answering by topic matching alone. Filler sessions contain unrelated conversation and extend the history to the target length. Every item begins with a structured plan specifying the question, gold answer, and required evidence before dialogue or audio is written.
-
Text and audio sources. Evidence and haystack dialogues are written as text first. Filler text is taken from InstructS2S and trimmed to match VoxMem session lengths. User turns are synthesized with Higgs-TTS-3, using a fixed VCTK voice per user. Paralinguistic cues are introduced through style controls, and environmental sounds from ESC-50 are mixed in at 10dB SNR. Questions are rendered into natural language using Gemini-3.7-Flash and GPT-5.6-Luna. In the final input, user turns are audio, assistant turns are text, and session boundaries and timestamps are included.
-
Scale. The construction pipeline generates more than 30,000 plans. The final answerable set contains 669 questions, with answer-refusal variants derived from existing answerable questions. Each answerable question is embedded in histories at four lengths: 8K, 16K, 32K, and 64K tokens measured with a Whisper encoder, approximately 2.5 to 20 minutes of audio. Longer histories strictly extend shorter ones, and evidence and haystack sessions are distributed uniformly except where temporal order matters for temporal evolution tracking.
-
Filtering and quality control. Session-level checks verify transcript fidelity, speaker consistency, target-cue perceptibility, and deterministic consistency of paired paralinguistic and environmental renditions. Question-family checks verify that evidence uniquely determines the answer, the question requires the intended operation, audio-native questions cannot be answered from the transcript, and full histories contain no leaked or alternative answers. Items that fail are revised and rechecked or discarded. A text-only classifier also cannot reliably distinguish evidence from haystack sessions, confirming that surface patterns do not reveal the answer.
-
Usage. VoxMem is an evaluation-only benchmark, not a training set. It is designed to evaluate LALMs that accept audio input but do not generate speech. Every model receives the same history with audio user turns and text assistant turns, and performance is compared across the four context lengths so that changes can be attributed to growing context rather than changing evidence.
Method
The authors construct the VoxMem benchmark along a two-dimensional taxonomy, utilizing a multi-session conversation structure composed of three distinct session types to ensure both naturalism and controlled evaluation scenarios. Evidence sessions contain the specific information required to answer a query. Haystack sessions introduce plausible but misleading content on the same topic, forcing the model to reason over the full acoustic and semantic context rather than relying on topic matching. Filler sessions consist of ordinary conversations unrelated to the question, serving to extend the history to the target context length.
As shown in the figure below, this structure is applied across different memory types. For instance, in an Information Extraction scenario involving speaker identity, the evidence session contains the answer spoken by the querying user, while a haystack session features a different user discussing the same topic with different content.
The construction of every item follows a rigorous pipeline, as illustrated in the framework diagram.
The process begins with Task and Evidence Design, where a structured plan specifies the question, gold answer, and required evidence. Evidence is distributed across sessions based on the operation type: a single session for Information Extraction, multiple for Multi-Session Reasoning, and an ordered sequence for Temporal Evolution Tracking. During the Dialogue and Audio Realization phase, evidence sessions are written as user-assistant dialogues. To prevent answer leakage, the two sides are authored independently. User turns are synthesized using a TTS system with fixed voices, and paralinguistic cues or environmental sounds are introduced.
In the Variants and Context Assembly stage, the authors create Answerable and Answer Refusal variants. Haystack and filler sessions are added as distractors. Histories are assembled at four lengths (8K, 16K, 32K, and 64K tokens) by strictly extending shorter histories with more distractors, ensuring the question and evidence remain identical. Finally, the Validation and Final Benchmark stage involves multi-stage quality control.
The authors implement a two-level quality control process. Session-level validation targets failures within a single session, verifying transcript fidelity, speaker consistency, and the perceptibility of acoustic cues. Question-family validation targets failures emerging in the full history. This includes checking answer and operation validity, ensuring acoustic necessity by rejecting questions answerable from text alone, and verifying full-history validity to ensure no leaked or alternative answers exist. Items failing these checks are revised and rechecked.
Experiment
The VoxMem benchmark evaluates spoken conversational memory across a two-dimensional taxonomy of acoustic evidence and memory operations, testing 15 large audio-language models on histories from 8K to 64K tokens with LLM-judged short answers. The experiments show that current models remain unreliable, with no model exceeding 40% overall accuracy at 32K, and that non-lexical acoustic information such as speaker identity, paralinguistic cues, and environmental sounds is substantially harder to retain than speech semantics. Memory operation difficulty depends on the evidence type, while refusal patterns indicate that models often abstain because of weak acoustic processing rather than genuine awareness of missing evidence. Longer histories reduce access to the same evidence, with speaker and environmental memory degrading fastest and error profiles differing across evidence types and operations.
The comparison shows that existing spoken-memory benchmarks cover acoustic evidence and memory operations only partially. Most emphasize linguistic content or selected acoustic cues and evaluate only retrieval or integration, while temporal evolution tracking and answer refusal are absent. The benchmarks also rely on single-session monologues or dialogues, with controlled scaling and evidence provenance rarely reported. Prior benchmarks rarely test speaker identity, paralinguistic cues, or environmental sound; several focus only on lexical content. Memory operation coverage is narrow: all listed benchmarks include information extraction, a few add multi-session reasoning, and none include temporal evolution tracking or answer refusal. All listed benchmarks use single-session histories, with no multi-session evaluation. Evidence provenance is missing for most listed benchmarks, with only one dialogue benchmark providing it.
The evaluation taxonomy is uneven, with paralinguistic and semantic questions dominating and answer refusal less common. Experimental results show that longer histories reduce access to previously available evidence, especially for speaker and environmental sound memory, while paralinguistic understanding remains low at all context lengths. Error attribution also reveals distinct failure modes across operations and evidence types, such as binding errors for speaker information and localization errors for temporal evolution. The taxonomy is unevenly distributed: paralinguistic and semantic questions are the most common evidence categories, with answer refusal less frequent across all operations. As history grows, models retain less access to evidence, and speaker identity and environmental memory degrade fastest; paralinguistic cues remain difficult at every context length.
The benchmark includes several hundred questions and several thousand instances drawn from tens of thousands of sessions, with a smaller subset of evidence sessions. Sessions are relatively compact, averaging about nine turns and a few clips per session. Controlled context expansion from 8K to 64K shows accuracy declines as history grows, with the steepest relative drops for speaker identity and environmental sound. The dataset contains 799 questions and 3,196 instances, but only a small fraction of sessions serve as evidence sessions. Longer context windows reduce accuracy across evidence types, with speaker and environmental memory degrading faster than speech semantics and paralinguistic cues.
For Gemini-3.7-Flash on answerable acoustic-memory questions, replacing full audio with transcript-only turns causes a large overall accuracy drop, confirming reliance on non-text acoustic cues. The impact is highly uneven: speech semantics remains relatively stable, while speaker identity, paralinguistic cues, and environmental sounds lose most of their accuracy. Overall full-audio accuracy falls from about 55 percent to below a quarter with transcript-only input. Speech semantics is least affected, while speaker identity drops roughly 60 points to just over 10 percent, and paralinguistic and environmental scores fall near zero.
Spoken-memory accuracy declines as reference history grows, with both proprietary and open-weight models losing accuracy between 8K and 32K. The drop is steeper for speaker identity and environmental sound than for speech semantics and paralinguistic cues. Paralinguistic cues show the lowest absolute accuracy but decline at a rate similar to speech semantics, indicating that baseline difficulty and history sensitivity are distinct failure modes. At 64K, speech semantics and paralinguistic cues retain about 70% of their 8K accuracy, while speaker identity and environmental sound retain about 63% to 67%. Proprietary models score above open-weight models at 8K and 32K, but both groups lose accuracy as history grows. Speaker errors are most often binding failures, whereas paralinguistic errors are dominated by localization failures and environmental errors split between localization and unsupported answers.
Existing spoken-memory benchmarks cover acoustic evidence and memory operations only partially, often missing speaker identity, paralinguistic and environmental cues, temporal evolution tracking, answer refusal, multi-session histories, and evidence provenance. The benchmark's controlled context expansion shows that both proprietary and open-weight models lose access to earlier evidence as history grows, with speaker identity and environmental sound degrading more steeply than speech semantics, while paralinguistic cues remain difficult across all context lengths. Replacing audio with transcripts causes a large overall drop and leaves speaker, paralinguistic, and environmental performance near zero, confirming that models rely on acoustic cues beyond text. Error attribution reveals distinct failure modes, including binding errors for speaker information and localization errors for paralinguistic and temporal evolution questions.