The Science Behind DAP™: Personalized Audio That Keeps Up with Real Life

Beyond “Movie Mode”
Most audio devices offer presets - Movie Mode, Rock Mode - that apply one fixed filter to everything. They don’t know whether you are hearing dialogue or a drum solo, and they certainly don’t know how you hear. For listeners with hearing loss, who may need the speech clearer, the background quieter, or the talking slower, a fixed preset is not personalization. It’s a label.
At Forum Acusticum 2025, the 11th Convention of the European Acoustics Association, our team at Bettear presented the research framework behind DAP™ - Deep Audio Processing - showing how true personalization works, in both real-time and offline listening.
How It Works

DAP™ combines classical audio engineering with modern machine learning in one pipeline:
- Separate. A deep neural network splits the incoming audio into speech, music and sound effects - so each can be treated on its own.
- Understand. A classifier continuously recognizes what is happening in the audio - conversation, music, ambient noise - and a speech-rate estimator measures how fast the speaker is talking.
- Personalize. The system then rebuilds the mix to the listener’s own targets: lifting speech to their preferred level above the background, and gently slowing speech that is too fast to follow.
The hardest part is time. Slow the speech down and the audio drifts out of sync with the video - lips stop matching words. DAP™ solves this with non-uniform time-scale modification: it stretches only the fast speech, compresses pauses and silences to win the time back, and keeps the whole signal within a strict synchronization budget - in the published example, a maximum deviation of 240 milliseconds. The result: slower, clearer speech that still matches the picture.
Grounded in Listener Evidence
How slow is slow enough? The paper includes a study with 44 listeners with bilateral hearing loss, who adjusted a narrator’s speech rate in background noise to the fastest rate they could still understand. The comfortable maximum fell from 10.94 phonemes per second for mild hearing loss to 8.51 for profound hearing loss - evidence that speech-rate needs, like mixing needs, differ by hearing profile. These measured preferences become the default presets DAP™ starts from.
One Framework, Many Experiences
Because the same components adapt to the content and the listener, the framework serves live streaming, cinema, teleconferencing and music playback alike - enhancing intelligibility where it is needed while preserving the character of the original sound. It is the engineering foundation of what DAP™ delivers today: audio produced once, made personal for everyone.
The full peer-reviewed paper - Bar-Yosef, Thor & Fink (2025), “The Revolution is Here: Deep Audio Processing Redefines Offline and Real-Time Audio Experiences” - is available via the European Acoustics Association.

