Spatial Audio Accessibility Demo

Move the player toward the target using arrow keys or WASD. Distance sets loudness, left/right sets panning and inter-aural time difference (ITD), and up/down sets a low-pass filter. Each cue can be toggled on or off.

Controls

Audio stopped

Spatial cues
ITD tuning 1.00×

World

Live values

Distance
0.0
dx (right +)
0.0
dy (up +)
0.0
Gain
0.0 dB
Pan
0.00
ITD
0 us
Cutoff
0 Hz

Keyboard help

Click or tab to the world, then:

How it works

This is a small demo of a spatial-audio model built for accessibility. It was inspired by an accessibility-modding discussion: a player of a 2D game could not judge distance to enemies from sound alone and asked for logarithmic distance dropoff. The fix decomposes localization into three independent axes, because the ear judges each direction from different information. Each axis drives a different psychoacoustic cue, and each cue can be toggled on and off so you can hear exactly what it adds.

Distance → loudness (logarithmic)

How far away the target is sets its overall loudness. Loudness perception is roughly logarithmic, which is why we measure sound in decibels. A naive linear falloff crams almost all the audible change into the last stretch near maximum range, leaving the middle of the range flat and unjudgeable. A decibel curve makes equal ratios of distance feel like equal loudness steps, so the whole range stays usable.

Left–right → panning + time delay

Horizontal position is encoded two ways at once. Stereo panning (an inter-aural level difference) on its own is a fairly weak and sometimes ambiguous side cue. Adding an inter-aural time difference (ITD) — the tiny delay before a sound reaches the far ear — makes left–right placement much more solid. Together, level and timing are the classic duplex theory of localization. The real ITD is only about 0.66 ms at most, and with no full head-related filtering behind it that can be hard to hear in isolation, so an exaggeration factor (the ITD tuning slider, or [ / ]) lets you scale it up until the timing cue is obvious — 1× is physically accurate.

Up–down → low-pass filter (brightness)

Panning carries nothing vertical, so up/down needs its own cue. Real elevation perception relies partly on how the outer ear filters sound: lower sources tend to lose high-frequency energy and sound duller. Tying a low-pass filter to vertical position gives a "bright = up, muffled = down" axis. In a game this doubles as a rough occlusion cue — something below you is often literally underground.

Altogether it is a lightweight, hand-rolled approximation of binaural spatialization built from cheap DSP primitives, with every cue exposed as an independently tunable knob. That decomposition is the point, and it has a practical benefit for accessibility: each cue can be exaggerated beyond physical realism for clarity. Every computed value is shown in the live readout and announced to your screen reader.

For the full psychoacoustic model, the math, and the code, see the GitHub repository.