Ambi Encoder Vox

s3g Ambi Encoder Vox is a 16-voice voicebank and WORLD-based vocal instrument with a fixed 64-channel, first-through-seventh-order ACN/SN3D output. It maps a typed phrase across independently pitched and positioned vocal sources; the 64 channels describe the ambisonic field, not the source-voice count.

WORLD analysis separates a recorded voice into pitch, spectral envelope, and aperiodic or noise components. Ambi Vox resynthesizes those components at new pitches and timings while retaining more of the source's vocal character than ordinary sample-rate transposition.

Ambi Encoder Vox with phrase source, voicebank and WORLD controls, ensemble field, motion, and output.

Workflow

  1. Insert the instrument on a 64-channel REAPER track, choose an initial factory preset, and select the ambisonic ORDER.
  2. Follow it with an ambisonic decoder.
  3. Choose FREE, MIDI, or BOTH.
  4. Load a UTAU-style voicebank folder or a vocal WAV file.
  5. Enter a phrase, then set 1-16 voices, an ensemble direction, pitch relationships, and voice-delay spread.

Presets

The title-bar PRESET menu provides thirteen starting points spanning solo speech, unison, SATB and double choirs, chorus, round, cluster, spectral color, texture, spatial orbit, and a full 16-voice choir. Every factory preset starts at 3OA and sets Output to -6 dB. Selecting one updates the instrument controls and phrase with a short output transition; ORDER can then be changed manually from 1OA through 7OA.

LOAD and SAVE beside the preset menu read and write .s3gvox JSON files. A user preset stores the editable instrument parameters, lyric cue sheet, phrase-generator recipe, and camera view. Vocal WAV files, voicebank folders, and analyzed WORLD audio are deliberately not embedded: loading a preset keeps the currently loaded sound source in place, so one performance design can be auditioned with different voices. The normal CLAP project state also remembers the selected factory or user preset name.

Phrase Source

The LOAD button accepts either a voicebank folder containing WAV files and oto.ini, or a single vocal WAV file. Every loaded voicebank alias is analyzed and resynthesized by WORLD before playback. Analysis follows pitch with Harvest, DIO, and StoneMask, captures the spectral envelope with CheapTrick, and measures aperiodic energy with D4C.

s3g Vox Builder prepares compatible voicebank folders from a continuous vocal recording. It detects the ordered aliases, shows editable segment boundaries, estimates root pitch and voicing with WORLD, auditions each slice, and exports WAV files with oto.ini timing.

The first voicebank load writes a persistent analysis cache in ~/Library/Caches/org.s3g.s3g-dsp/AmbiVoxWorld. Later loads reuse it until a source WAV changes.

Voicebank preparation caches five formant-preserving WORLD pitch anchors from two octaves below through two octaves above each alias's detected base pitch. Runtime voices select the nearest anchor, which keeps wide MIDI and scale intervals closer to the original vocal character.

Words are mapped to available voicebank aliases and romaji sample keys by a small English grapheme-to-phoneme layer. The phrase panel shows the resolved alias sequence and a live voice-1 timeline. SPEAK plays every resolved alias once, reads the complete phrase in order, then loops after the phrase-end pause. SING can sustain vowels while keeping every voice on the same phrase sequence. TEXTURE treats loaded samples as a continuously moving source field.

PHR SPREAD lets active voices read different parts of the source phrase. At 0% every voice begins on the same event; at 100% voices are distributed across the complete phrase. Voicebanks always move by complete alias or rest events, so an offset never begins halfway through a syllable. A single loaded WAV uses the same control to distribute starting positions across its loop, while the internal voice model offsets compiled phoneme events.

VOICE STEP delays each successive rendered voice by as much as 1000 ms; VOICE DEV adds a repeatable deviation around that cascade. These are short entrance and timing offsets rather than phrase-position controls. They are smoothed and retarget dual delay taps with a 120 ms crossfade. Moving either control does not clear a delay, recompile a bank, or restart the phrase.

Lyrics Field

The left-side LYRICS page is a cue sheet. Each non-empty line becomes one independently compiled lyric cue, up to 32 cues. The cue map shows the selected line and the line currently being performed. Clicking a cue selects it; the single-line phrase field remains a quick editor for that selected cue. PHR SPREAD is always contained within the active cue, so a round can begin at different syllables without spilling into the next line.

LYRICS view with a populated multi-cue sheet and its compiled cue map.

TEXT opens the cue editor. GENERATE opens a deterministic pseudo-language composer whose output returns to that same editable cue sheet. FORM chooses drift, call-and-response, round, bloom, or pulse relationships; COLOR biases vowel families; and GROUPS, CUES / GROUP, LENGTH, VARIATION, HARDNESS, REST, and SEED shape the result. The available syllables are checked against the loaded voicebank. GENERATE restarts from the seed, MUTATE keeps the underlying group motifs while deriving a new generation, and APPEND adds a generation without exceeding the 32-cue limit.

MIDI CH defaults to channel 16. In either MIDI cue mode, notes on that channel are consumed as cue control and do not trigger vocal voices. Other MIDI channels retain their normal instrument behavior. CUE BEATS sets the window length for TRANSPORT; clicking a line while AUTO is active restarts that cue's single performance. Cue selection occurs at an audio processing boundary, and voice timing continues to control performance inside the selected line.

Voicebank Timing and Pronunciation

The renderer honors each oto.ini entry's offset, consonant/fixed region, cutoff, preutterance, and overlap. Adjacent aliases crossfade using their preutterance and overlap values, and BLEND scales that transition. In SPEAK, the complete alias is one-shot and never wraps inside its phrase event. In SING, the consonant region runs once and the vowel can sustain after the fixed region. For CV banks with zeros in every timing field, Ambi Vox derives a short onset and vowel region from the recording.

A bank can override the built-in word mapping with an optional s3g-pronunciations.txt file in its root folder. Put one lower-case word and its alias sequence on each line:

fire=fa i ya
voices=bo ka ru
sorrow=so ro

Aliases may be separated by spaces, commas, semicolons, or vertical bars. Lines beginning with # are ignored. Every alias on a line must exist in the loaded bank for that override to be used.

Vocal Layer

WORLD Voice Model

These controls use the analysis attached to each loaded WORLD source. They remain neutral when centered and are smoothed before reaching the phase-vocoder resynthesis path.

Ensemble Direction

SPREAD contracts or expands an ensemble's interval pattern. In MIDI, held notes always become separately allocated spatial voices, independent of the ensemble choice. In BOTH, a non-INDIVIDUAL ensemble can still take the newest held note as the common root of the free-running choir.

The field keeps each source point's azimuth/elevation/distance color and adds a muted ensemble-role frame. Chorale points are marked S, A, T, or B; chorus and round cohorts receive corresponding labels and faint same-role links. The outer square pulses at each voice's phrase position, making staggered entries visible without confusing orchestration with spatial location. See Interpreting Color for how these two color layers differ.

MIDI and Phrase Voices

In FREE, active points take scale degrees or the selected ensemble relationship around the root. In MIDI, the first held note activates point 1, the second activates point 2, and so on through the selected VOICES limit. Releasing a note silences and frees its point; the next note reuses the lowest free point, and excess notes steal the oldest active point. The field display mutes inactive squares so this allocation remains visible. BOTH retains the free-running field and lets held MIDI notes influence its pitch behavior.

Playback Timing

The playback panel changes with VOICE mode so that only relevant timing controls are shown. In SPEAK and SING, PACE schedules phrase events; word spaces and punctuation produce explicit short, medium, and sentence-length gaps. In TEXTURE, SPEED, loop start/end, freeze, and position control the continuously read source. Texture speed has no effect on phrase playback.

STRETCH changes playback duration from 0.25x to 4x without changing the selected musical pitch. In SPEAK, each alias remains one-shot; in SING, the fixed consonant region is preserved while vowels can sustain. TRANSIENT spans soft to preserved attacks in every voice mode: lower settings use fewer phase resets, longer event fades, reduced fixed-consonant energy, and per-voice onset conditioning; higher settings retain the source articulation and its potentially percussive character. BLEND sets adjacent phrase-alias overlap. For voicebanks, the phase vocoder handles residual tuning between cached WORLD pitch anchors rather than carrying the full pitch shift.

Output and Plosive Control

POP FILTER dynamically reduces short low-frequency pressure bursts and the broadband crack that can make a synthetic plosive resemble a drum hit. It first softens abrupt per-voice onset regions, then derives one shared detector from the strongest low-band and broadband onset across the active ambisonic bus. Matching reduction is applied to every output channel, preserving spatial coherence even when a directional burst is weak in W but strong in higher-order channels. Sustained bass and consonant detail return as the detector releases; 0% leaves this stage transparent and 100% provides the strongest plosive reduction.

Motion

MANUAL, ORBIT, FLOW, PATH, and PULSE move voices in the AED field. Motion SPEED changes only this spatial movement and does not alter phrase or audio playback speed. Motion can run freely or follow the REAPER transport.

Order and Routing

The output bus remains 64 channels wide even though synthesis is capped at 16 source voices. ORDER activates the corresponding ACN/SN3D channels and clears all higher outputs.

OrderActive channels
1OA4
2OA9
3OA16
4OA25
5OA36
6OA49
7OA64