How it works
This is a small demo of a spatial-audio model built for
accessibility. It was inspired by an accessibility-modding discussion: a
player of a 2D game could not judge distance to enemies from sound alone
and asked for logarithmic distance dropoff. The fix decomposes localization
into three independent axes, because the ear judges each direction from
different information. Each axis drives a different psychoacoustic cue, and
each cue can be toggled on and off so you can hear exactly what it adds.
Distance → loudness (logarithmic)
How far away the target is sets its overall loudness. Loudness perception is
roughly logarithmic, which is why we measure sound in decibels. A naive
linear falloff crams almost all the audible change into the last stretch
near maximum range, leaving the middle of the range flat and unjudgeable. A
decibel curve makes equal ratios of distance feel like equal
loudness steps, so the whole range stays usable.
Left–right → panning + time delay
Horizontal position is encoded two ways at once. Stereo panning (an
inter-aural level difference) on its own is a fairly weak and
sometimes ambiguous side cue. Adding an inter-aural time
difference (ITD) — the tiny delay before a sound reaches the far ear —
makes left–right placement much more solid. Together, level and timing are
the classic duplex theory of localization. The real ITD is only about
0.66 ms at most, and with no full head-related filtering behind it that can
be hard to hear in isolation, so an exaggeration factor (the ITD
tuning slider, or [ / ]) lets you scale it up until
the timing cue is obvious — 1× is physically accurate.
Up–down → low-pass filter (brightness)
Panning carries nothing vertical, so up/down needs its own cue. Real
elevation perception relies partly on how the outer ear filters sound:
lower sources tend to lose high-frequency energy and sound duller. Tying a
low-pass filter to vertical position gives a "bright = up, muffled = down"
axis. In a game this doubles as a rough occlusion cue — something below you
is often literally underground.
Altogether it is a lightweight, hand-rolled approximation of binaural
spatialization built from cheap DSP primitives, with every cue exposed as
an independently tunable knob. That decomposition is the point, and it has
a practical benefit for accessibility: each cue can be exaggerated beyond
physical realism for clarity. Every computed value is shown in the live
readout and announced to your screen reader.
For the full psychoacoustic model, the math, and the code, see the
GitHub
repository.