Advanced / Audio · guide
Procedural audio (createAudio)
Current Apexify.js 6.0.0 documentation for Procedural audio (createAudio).
painter.createAudio synthesizes 16-bit PCM WAV Buffers — presets, custom multi-layer sounds, timelines, and mixes. Output is ready for mixAudio, videoPipeline().audio(), disk via save, or any API that accepts a WAV buffer.
Phase 9 hardens this surface with pre-allocation resource checks, operation-local seeded randomness, stricter DSP validation, and RIFF-safe PCM16 handling. createAudio remains a complete-buffer API: it returns a finished WAV Buffer, not a streaming audio object.
Hub: Audio (advanced) · Runtime limits: Runtime validation & resource governance · Mux on video: Video audio operations · Pipeline tracks: Video pipeline
Entry points
| Method | Returns | Use when |
|---|---|---|
preset(name, overrides?) | Buffer | Built-in SFX (laser, coin, jump, …) |
synth(options) / custom(options) | Buffer | Full SynthSoundOptions (layers, ADSR, filters) |
sequence({ events }) | Buffer | Timed preset/custom events on one WAV |
compose({ clips }) | Buffer | Overlapping clips with pitch, pan, fades, tone |
mix(inputs, options?) | Buffer | Sum buffers at t = 0 or with per-input at |
save(wav, filePath) | Promise<void> | Write WAV to disk |
listPresets() | SynthPresetInfo[] | Names, descriptions, default durations |
presetNames | readonly array | All SynthPresetName values |
Types export from apexify.js/types: SynthPresetName, SynthSoundOptions, SynthLayer, SynthSequenceOptions, SynthComposeOptions, …
Deterministic seeded audio
Noise-based synthesis is nondeterministic by default. Supply seed when you need byte-stable procedural output for tests, cache keys, reproducible renders, or parallel jobs.
seed accepts a safe integer or a 1–256 character string on sound, sequence, compose, clip, and mix options where the type exposes it.
Seeded layers/events/clips derive independent random streams. Concurrent seeded renders do not share RNG or filter state. Omitting seed intentionally keeps stochastic sounds nondeterministic.
Presets
39 built-in names (game/UI oriented): laser, explosion, coin, jump, whoosh, menuSelect, footstep, thunder, …
SynthPresetOverrides: volume, transpose, plus any SynthSoundOptions field layered on the preset definition. Preset definitions are cloned before use, so caller mutation does not alter the shared catalog.
Custom sounds — synth / custom
Each SynthLayer can use waveforms sine, square, sawtooth, triangle, noise, pink:
| Field | Role |
|---|---|
frequency / frequencyEnd | Start Hz and optional sweep for tonal waveforms |
duration | Layer length (seconds) |
delay | Offset before layer starts |
gain, detune, pan | Level, cents, stereo pan |
adsr | Attack / decay / sustain / release |
vibrato, tremolo | LFO depth + rate |
filter | lowpass / highpass + cutoff + optional Q |
noiseMix, partials | Noise blend and harmonic ratios on tonal layers |
sampleRate (default 44100), channels (1 or 2), limiter (default true) apply at the sound level.
Noise/pink layers reject tonal-only controls such as frequency sweeps, detune, vibrato, and harmonic partials instead of silently ignoring them. Filter cutoff must be below the selected sample rate's Nyquist frequency.
Timelines — sequence
Events fire at at (seconds). Each event uses exactly one of preset or options (SynthSoundOptions), plus optional gain.
tail pads silence after the last event ends. Overlapping events sum in the same buffer (limiter applied). The renderer keeps the final output plus one event working buffer instead of retaining every rendered event in memory.
Overlapping clips — compose
compose is for dense beds: multiple presets, custom sounds, or existing WAV buffers on one timeline with per-clip shaping.
SynthComposeClip field | Effect |
|---|---|
at, duration, sourceStart | Placement and trim |
preset / sound / wav | Source (exactly one required) |
gain, volume, transpose, detune, pitch, speed | Level and pitch/playback-rate controls |
pan, fadeIn, fadeOut | Stereo and edges |
noise, filter, quality | Post-tone on the clip |
overrides | Preset tweaks |
seed | Deterministic clip-local stochastic processing |
Top-level SynthComposeOptions: duration, tail, masterGain, postHighpassHz (rumble cleanup), noiseGateThreshold (quiet overlap hiss), and optional seed.
For synthesized preset/custom sources, transpose, detune, and pitch modify oscillator pitch. Existing WAV clips do not accept those procedural pitch controls; use speed instead. speed is playback-rate resampling: it changes duration and pitch together and is not pitch-preserving time stretch.
The composer keeps the final output plus one processed clip working buffer rather than retaining every clip at once.
mix
Combine one or more Buffers (or inline preset/sound definitions where supported):
Use when you already have WAV buffers and need a single file without a full compose timeline. If any input uses timeline/clip controls, mix routes through the validated composition path.
Resource limits and allocation behavior
Audio requests are validated before large sample allocations. Defaults are controlled by the shared Apexify runtime configuration:
| Limit | Default |
|---|---|
maxAudioDurationSeconds | 600 seconds |
maxAudioSampleRate | 192,000 Hz |
maxAudioChannels | 2 |
maxAudioEvents | 20,000 |
maxAudioLayers | 1,024 |
maxAudioPartials | 4,096 |
maxAudioBytes | 256 MiB |
maxAudioBytes is a peak working-memory guard, not just a final WAV-size limit. Apexify.js accounts for the Float32 render buffer, decoded/resampled sources where applicable, transient event/clip buffers, and the final PCM16 WAV allocation when those buffers coexist.
Lower limits for multi-tenant or memory-constrained services. Do not raise them merely to suppress ApexifyResourceLimitError.
Because the API returns a complete WAV Buffer, Phase 9 deliberately enforces bounded complete-buffer rendering rather than claiming streaming synthesis.
WAV behavior
Generated output is mono/stereo PCM16 RIFF/WAVE. Internally, WAV inspection/decoding validates chunk boundaries and does not assume fmt and data occur at fixed offsets. Unsupported formats, bit depths, malformed chunk sizes, incomplete frames, inconsistent byte-rate/block-alignment metadata, and oversized decoded audio are rejected before unsafe allocation.
Put audio on video
Single operation — mixAudio
Pass any createAudio Buffer as MixAudioOverlayClip.source:
keepOriginalAudio: false → overlays only (no silent bed under SFX). Full field reference: Video audio operations.
Editor stack — videoPipeline().audio()
Declare tracks without pre-building every buffer:
type | Fields |
|---|---|
preset | preset, startTime, gain, volume, transpose |
synth | sound, startTime, gain |
sequence | events, startTime, tail, masterGain |
wav | Pre-built Buffer, startTime, volume |
file | Path, URL, or buffer (external assets) |
See Video pipeline — Audio layers.
Standalone export
vs FFmpeg video audio
createAudio | createVideo audio keys | |
|---|---|---|
| Input | Generated PCM | Existing video/audio files |
| Output | WAV Buffer | MP4 (or extracted audio file) |
| FFmpeg | Only when muxing to video | Always for video ops |
| Typical use | Game SFX, UI sounds, bots | Podcast normalize, ducking, strip track |
Tips
- Add
seedwhen output must be reproducible across repeated or concurrent renders. listPresets()before hard-coding names in editor UIs.- Prefer
sequencefor linear SFX chains;composewhen clips overlap with independent pitch/fades. - For long gameplay videos, generate SFX once, then
mixAudioor pipelinetype: 'wav'— avoid re-synthesizing on every export pass. - Keep deployment-specific audio limits below what the Node process can safely hold at peak.
- Pair with GIF /
animateby muxing SFX after frame encode.