Contents
Table of Contents
01 / How it works
Keep the speech codec. Add spatial tokens.
The pretrained mono codec carries speech content through the omnidirectional W
channel.
FOAdapter encodes spatial information in a separate low-rate
stream and reconstructs the complete FOA signal.
Compact spatial representationTwo 2,048-entry RVQ codebooks at 12.5 tokens/s add just 275 bit/s.
One adapter, multiple codecsThe same adapter is evaluated with DAC, Mimi, and MOSS without retraining.
Spatially guided learningSpatial supervision and contrastive learning encourage separation of source direction from speech content.
Headline results are from the Spatial LibriSpeech test set in the paper. Moving-speaker and speech-LM examples are supplementary demonstrations.
02 / Listen & compare
FOA speech reconstruction
Compare the original sound field with FOAdapter, channel-independent (CI) codecs, and Opus mapping family
2 using the same speech excerpt. The top-down view shows the target source azimuth relative to a listener facing
front.
FOAdapter provides highly
bitrate-efficient spatial speech reconstruction, adding only 275 bit/s of spatial
information for a total bitrate of 1.275 kbps with MOSS-8.
Headphones recommended. We binaurally rendered all FOA samples on this page using the HRTFs from the study below. Listen for speech quality and the perceived direction of the speaker.
Makoto Otani, Tatsuya Hirahara, and Shiro Ise, “Numerical study on source-distance dependency of head-related transfer functions,” The Journal of the Acoustical Society of America, vol. 125, pp. 3253–3261, 2009.
In neural codec names, the number after the hyphen indicates the number of codebooks used (e.g., MOSS-8 uses 8 codebooks). CI encodes each FOA channel independently. Rates are total codec bitrates, including FOAdapter’s spatial stream. A shared playback gain is used for all systems within each scene.
03 / Swap the spatial tokens
Same speech, a different direction
Keep the content tokens from source A and
replace
A's spatial tokens with those from source B.
The frozen FOAdapter combines A’s MOSS-8 speech
reconstruction with B’s spatial tokens.
Listen for A’s words
shifting toward B’s direction.
Content tokens ASpatial tokens BSwapped speech
Listen to the two original sources, then the result: A’s speech content combined with B’s spatial information.
Compasses show source azimuths and intended output directions, not measured output directions. All three clips in each example share one playback gain.
04 / Out-of-distribution test
OOD: a real-world moving speaker
This example tests out-of-distribution (OOD) generalization: the adapter was trained on simulated speech with fixed source and receiver positions, then applied to a real STARSS23 recording with a moving speaker.
Spatial LibriSpeech
STARSS23 · Same adapter, no retraining
Reference and reconstruction share the same playback gain. STARSS23 dataset ↗ · License
05 / From tokens to spatial speech
Speech language model integration
We fine-tuned MOSS-TTS with LoRA to generate spatialized speech in a given direction. A direction-conditioned spatial prediction branch produces spatial tokens alongside speech tokens, and the decoded speech and spatial tokens are combined by the frozen FOAdapter.
Input text
Each compass shows the requested fixed direction. All five variants share the same speech tokens and playback gain.