FOAdapterA Plug-and-Play Spatial Tokenizer for Low-Bitrate FOA Speech Coding

Kirak Kim1*· Yoonjeong Park2*· Juhan Nam1,2· Sungyoung Kim1· Minje Kim3

1 Graduate School of Culture Technology, KAIST2 Kim Jaechul Graduate School of AI, KAIST

3 Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign

* Equal contribution

FOAdapter enables low-bitrate first-order ambisonic (FOA) speech coding while preserving both linguistic information
and spatial cues. It extends a frozen pretrained monaural speech codec with a separate spatial token stream.

275 bit/sSpatial token stream
1.275 kbpsTotal with MOSS-8
6.68°MOSS-8 + FOAdapter DOA error

01 / How it works

Keep the speech codec. Add spatial tokens.

The pretrained mono codec carries speech content through the omnidirectional W channel.
FOAdapter encodes spatial information in a separate low-rate stream and reconstructs the complete FOA signal.

FOAdapter architecture. The FOA input’s W channel passes through a frozen pretrained monaural encoder, quantizer, and decoder to recover speech content. A trainable spatial encoder and two-stage RVQ produce spatial tokens at 12.5 tokens per second. The spatial decoder combines these tokens with W during training or decoded W at inference to reconstruct all four FOA channels.
Open full-size figure

Compact spatial representationTwo 2,048-entry RVQ codebooks at 12.5 tokens/s add just 275 bit/s.

One adapter, multiple codecsThe same adapter is evaluated with DAC, Mimi, and MOSS without retraining.

Spatially guided learningSpatial supervision and contrastive learning encourage separation of source direction from speech content.

Headline results are from the Spatial LibriSpeech test set in the paper. Moving-speaker and speech-LM examples are supplementary demonstrations.

02 / Listen & compare

FOA speech reconstruction

Spatial LibriSpeech

Compare the original sound field with FOAdapter, channel-independent (CI) codecs, and Opus mapping family 2 using the same speech excerpt. The top-down view shows the target source azimuth relative to a listener facing front.
FOAdapter provides highly bitrate-efficient spatial speech reconstruction, adding only 275 bit/s of spatial information for a total bitrate of 1.275 kbps with MOSS-8.

Headphones recommended. We binaurally rendered all FOA samples on this page using the HRTFs from the study below. Listen for speech quality and the perceived direction of the speaker.

Makoto Otani, Tatsuya Hirahara, and Shiro Ise, “Numerical study on source-distance dependency of head-related transfer functions,” The Journal of the Acoustical Society of America, vol. 125, pp. 3253–3261, 2009.

4 examples · 8-second excerpts

In neural codec names, the number after the hyphen indicates the number of codebooks used (e.g., MOSS-8 uses 8 codebooks). CI encodes each FOA channel independently. Rates are total codec bitrates, including FOAdapter’s spatial stream. A shared playback gain is used for all systems within each scene.

03 / Swap the spatial tokens

Same speech, a different direction

Spatial token swap

Keep the content tokens from source A and replace A's spatial tokens with those from source B.
The frozen FOAdapter combines A’s MOSS-8 speech reconstruction with B’s spatial tokens.
Listen for A’s words shifting toward B’s direction.

Content tokens ASpatial tokens BSwapped speech

Listen to the two original sources, then the result: A’s speech content combined with B’s spatial information.

Compasses show source azimuths and intended output directions, not measured output directions. All three clips in each example share one playback gain.

04 / Out-of-distribution test

OOD: a real-world moving speaker

OOD · Unseen recording conditions

This example tests out-of-distribution (OOD) generalization: the adapter was trained on simulated speech with fixed source and receiver positions, then applied to a real STARSS23 recording with a moving speaker.

Training settingSimulated rooms · Fixed positions

Spatial LibriSpeech

OOD evaluationReal recording · Moving speaker

STARSS23 · Same adapter, no retraining

Reference and reconstruction share the same playback gain. STARSS23 dataset ↗ · License

05 / From tokens to spatial speech

Speech language model integration

MOSS-TTS + LoRA

We fine-tuned MOSS-TTS with LoRA to generate spatialized speech in a given direction. A direction-conditioned spatial prediction branch produces spatial tokens alongside speech tokens, and the decoded speech and spatial tokens are combined by the frozen FOAdapter.

Input text

One utterance, five requested directions. The LoRA-fine-tuned TTS model generates a shared utterance; its spatial branch predicts a separate set of spatial tokens for each fixed direction: −90°, −45°, 0°, +45°, and +90°. Positive angles point left and negative angles point right; 0° is directly in front.

Each compass shows the requested fixed direction. All five variants share the same speech tokens and playback gain.