Audio Embodied Intelligence

Survey project page

Audio Embodied Intelligence

Auditory Perception, Reasoning, and Interaction for Multimodal Agents

Wentao Lei*, Tianxin Xie*, Kai Jiang*, and Li Liu

A unified survey of how sound helps embodied agents perceive hidden physical events, ground task intent, reason about users, and close the loop through action and interaction.

Definition of embodied audio and its roles in an embodied loop
3 auditory embodied loops
3 Percept, Reason, Interact stages
9 loop-stage categories

Core idea

Sound matters when it changes what an agent believes, does, or communicates.

The paper defines embodied audio as auditory input or output that is explicitly connected to a situated agent, environment, task, action, interaction, or evaluation. The key question is not only whether an audio model is accurate, but whether audition changes downstream state estimation, decision-making, action selection, or social response.

Physical Interaction Loop

Sound reveals spatial layout, contact dynamics, object state, materials, and action consequences.

Semantic Grounding-and-Action Loop

Speech, sound, and language convey task intent, referents, constraints, memory, and action choices.

Unified framework for embodied audio across Percept, Reason, and Interact stages
The proposed unified framework: three auditory roles cross the Percept, Reason, and Interact stages.

Taxonomy and scope

Auditory embodied loops as the unit of analysis

A work is included when sound is not merely a signal to be recognized, but part of a situated loop that connects auditory input or output to embodied perception, grounded reasoning, control, interaction, or evaluation.

Stage axis

Percept extracts reliable auditory or multimodal cues from situated observations. Reason transforms these cues into task-relevant state, grounded meaning, plans, user models, or social interpretations. Interact uses the inferred state or meaning to produce physical action, spoken response, expressive behavior, feedback, or recovery.

Loop axis

The loop axis specifies what the internal state represents: physical state in the Physical Interaction Loop, grounded task meaning in the Semantic Grounding-and-Action Loop, and user or social state in the Socio-Affective Interaction Loop.

z_t = f_percept(a_<=t, v_<=t, l_<=t, h_<=t) s_t = f_reason(z_<=t, x_<=t) (u_t, y_t) = f_interact(s_t, x_t)

In the LaTeX formulation, z_t denotes extracted auditory or multimodal cues, s_t denotes the task-relevant physical, semantic, or socio-affective state, u_t denotes physical action, and y_t denotes communicative or expressive output.

Crossing the three embodied loops with the three functional stages yields nine loop-stage categories:

Physical Interaction Loop process view
Physical Interaction Loop: acoustic cues support state estimation and feedback.
Semantic Grounding-and-Action Loop process view
Semantic loop: speech and sound connect task intent to grounded action.
Socio-Affective Interaction Loop process view
Socio-affective loop: vocal cues shape user state and expressive response.

Datasets, benchmarks, and metrics

From listening accuracy to behavioral consequence

Auditory embodied intelligence needs evaluation that isolates whether sound causally improves navigation, manipulation, dialogue, safety, or user experience. The paper organizes datasets, benchmarks, and metrics by the same three-stage pipeline.

Recommended ablations

Compare audio-aware agents against audio-off, corrupted-audio, delayed-audio, shuffled-audio, distractor-source, and oracle-audio variants. If behavior does not change, the system may merely contain audio input rather than use audition as embodied evidence.

Applications and deployment

Where audio-aware embodied agents matter

The survey groups applications by deployment need rather than by model family. In each setting, the value of sound is measured by the action or interaction it improves.