Physical Interaction Loop
Sound reveals spatial layout, contact dynamics, object state, materials, and action consequences.
Survey project page
Auditory Perception, Reasoning, and Interaction for Multimodal Agents
A unified survey of how sound helps embodied agents perceive hidden physical events, ground task intent, reason about users, and close the loop through action and interaction.
Core idea
The paper defines embodied audio as auditory input or output that is explicitly connected to a situated agent, environment, task, action, interaction, or evaluation. The key question is not only whether an audio model is accurate, but whether audition changes downstream state estimation, decision-making, action selection, or social response.
Sound reveals spatial layout, contact dynamics, object state, materials, and action consequences.
Speech, sound, and language convey task intent, referents, constraints, memory, and action choices.
Prosody and expressive cues support affect sensing, user modeling, trust, norms, and adaptive response.
Taxonomy and scope
A work is included when sound is not merely a signal to be recognized, but part of a situated loop that connects auditory input or output to embodied perception, grounded reasoning, control, interaction, or evaluation.
Percept extracts reliable auditory or multimodal cues from situated observations. Reason transforms these cues into task-relevant state, grounded meaning, plans, user models, or social interpretations. Interact uses the inferred state or meaning to produce physical action, spoken response, expressive behavior, feedback, or recovery.
The loop axis specifies what the internal state represents: physical state in the Physical Interaction Loop, grounded task meaning in the Semantic Grounding-and-Action Loop, and user or social state in the Socio-Affective Interaction Loop.
z_t = f_percept(a_<=t, v_<=t, l_<=t, h_<=t)
s_t = f_reason(z_<=t, x_<=t)
(u_t, y_t) = f_interact(s_t, x_t)
In the LaTeX formulation, z_t denotes extracted auditory or multimodal cues, s_t denotes the task-relevant physical, semantic, or socio-affective state, u_t denotes physical action, and y_t denotes communicative or expressive output.
Crossing the three embodied loops with the three functional stages yields nine loop-stage categories:
Datasets, benchmarks, and metrics
Auditory embodied intelligence needs evaluation that isolates whether sound causally improves navigation, manipulation, dialogue, safety, or user experience. The paper organizes datasets, benchmarks, and metrics by the same three-stage pipeline.
Compare audio-aware agents against audio-off, corrupted-audio, delayed-audio, shuffled-audio, distractor-source, and oracle-audio variants. If behavior does not change, the system may merely contain audio input rather than use audition as embodied evidence.
Applications and deployment
The survey groups applications by deployment need rather than by model family. In each setting, the value of sound is measured by the action or interaction it improves.