🔬 Applied Deep Tech R&D • Prototype benchmark v3.2 achieved: < 7.8ms edge pipeline latency. Technical Specs →
MisophoniaLabs Acoustic Neural Systems
Next-Generation Acoustic Neural Earwear

Silencing Biological Triggers.
Preserving Human Speech in Real-Time.

Misophonia Labs is developing wearable in-ear hardware powered by on-device neural filtering. Our pipeline surgically isolates chewing, breathing, and oral triggers within an 8-millisecond budget on embedded microcontrollers—eliminating distress while maintaining natural conversation.

Live Acoustic STFT Simulation
Window: 512 samples (32ms) | Hop: 128 (8ms) | Phase preserved
< 7.8 ms
End-to-End Latency

Well beneath the human auditory threshold for echo perception (10-15ms).

-28.4 dB
Selective Attenuation

Peak suppression depth for masticatory and salivary acoustic bursts.

3.92 PESQ
Voice Intelligibility

High perceptual speech quality score during active trigger elimination.

On-Device
Zero Cloud Reliance

Inference runs entirely offline on embedded micro-controllers.

The Neuro-Acoustic Problem

Why traditional Active Noise Cancellation is helpless against Misophonia

Misophonia is a recognized neurobehavioral condition where specific pattern-based biological sounds (chewing, swallowing, lip smacking, heavy breathing) trigger severe sympathetic nervous system arousal—causing involuntary panic, intense anger, and social withdrawal.

Commercial ANC headphones (like AirPods Pro or Sony WH-1000XM5) are designed strictly for stationary, predictable noise (airplane engines, air conditioning hums). They fail entirely with sudden, non-stationary oral transients.

Conversely, passive foam earplugs completely isolate individuals, creating severe conversational barriers during meals, lectures, and collaborative work.

Standard Active Noise Cancellation (ANC)

Inverts phase for stationary frequencies below 1 kHz. Cannot track or react to unpredictable, rapid biological transients such as mouth clicks or crunching.

Passive Attenuation (Earplugs)

Dampens speech frequencies uniformly across the spectrum. Makes interpersonal communication impossible, forcing individuals into social isolation.

The Misophonia Labs Selective Gate

Our Solution

Extracts real-time STFT spectral features, estimates a complex ratio mask, and surgically attenuates the trigger signature while maintaining speech harmonics and directional cues.

Signal Processing Core

Complex Ratio Masking (CRM) on Embedded DSP

Operating in the time-frequency domain to jointly reconstruct magnitude and phase with deterministic latency.

01

Dual-Domain STFT

Signals are framed at 16 kHz with a 512-point window (32ms) advanced by an 8ms hop size (128 samples). Both real and imaginary spectral coefficients are fed into the tensor core to preserve phase clarity.

N_FFT = 257 bins (514 Real/Imag features)
02

Deep Neural Mask Estimation

Trained on authentic trigger corpora blended with clean speech databases using an L1 Complex + Spectral Convergence composite loss function for rapid convergence.

Loss: L1_complex + Spectral_Conv + 0.5*L1_mag
03

Micro-Controller Quantization

The neural mask generator is quantized to 8-bit integer weights and tuned for sub-10ms execution budgets on high-speed embedded micro-controllers with hardware floating point units.

Inference cycle: < 3.6ms per frame

Targeted Trigger Suppression Benchmark (Tested across 150+ acoustic samples)

Mastication / Chewing
-28.2 dB
High attenuation; speech preserved
Lip Smacking / Saliva
-31.5 dB
Near-total transient suppression
Heavy Nasal Breathing
-24.0 dB
Adaptive comfort floor
Keyboard & Pen Clicking
-26.7 dB
Attenuates impulse spikes
Wearable Architecture

Designed for Inconspicuous, All-Day Wear

The hardware package pairs high-SNR front/rear MEMS microphones with an ergonomic custom in-ear acoustic canal. Dedicated low-power circuitry ensures rapid analog-to-digital conversion, neural computation, and balanced-armature output without perceptible phase delay.

1

Dual-Microphone Differential Acoustic Array

Captures directional sound vectors to preserve spatial localization so you always know where conversation partners are speaking from.

2

Hardware I2S Audio Bus & Low-Jitter Clock

Direct streaming interface bypassing operating system audio buffers, locking roundtrip audio processing latency below 8 ms.

3

Bluetooth 5.3 Low Energy Sync

Allows instant parameter updating from the companion mobile agent without routing live audio over wireless channels.

// System Technical Specifications Prot. v3.2
Sampling Rate: 16,000 Hz (16-bit uncompressed)
STFT Frame Length: 512 samples (32 ms window)
Hop Step Duration: 128 samples (8.0 ms budget)
Microcontroller Target: ARM Cortex-M7 (600 MHz) with DSP FPU
Mask Floor Limit: 0.05 (-26 dB floor; never mutes 100%)
Acoustic Transducer: Balanced Armature (Ultra-Low Distortion)
Target Battery Life: > 12 Hours continuous filtering
Model Context Protocol (MCP) Integration

Conversational Sensitivity Tuning

Every person with misophonia has a unique trigger signature. Instead of tedious manual equalization curves, our companion mobile app uses an intelligent conversational agent communicating via Model Context Protocol (MCP) to map user discomfort into runtime DSP filter parameters.

AI
Adaptive MCP Audio Agent
Tool-calling runtime synthesizer
U
"I am sitting in a quiet office meeting right now. A coworker next to me is crunching chips and whispering, and it's triggering panic. But I need to hear the speaker clearly."
AI
✓ MCP Tool Invoked: set_acoustic_profile • High-frequency transient filter boosted (-30 dB on 2.5–6 kHz crunch harmonics).
• Directional beamforming narrowed to front speaker (+4 dB vocal clarity).
• Masking floor lowered to prevent complete silence anxiety.

MCP-Driven Hardware Tool Calling

Users communicate in natural language with an intelligent conversational assistant. Through an MCP server interface, the agent directly invokes hardware tools to update the earphone's DSP registers and spectral coefficients over BLE.

Cognitive Behavioral (CBT) Exposure Habituation

Assists audiology professionals with gradual sound exposure therapy. The agentic assistant gently modulates attenuation levels over weeks, tracking emotional resilience without overwhelming the user.

Acoustic Context Logging & Privacy

Trigger detection statistics are stored locally on device and only metadata summaries are processed, maintaining strict privacy and ethical data standards.

Research Origins & Mission

Born out of applied acoustic engineering research in BahĂ­a Blanca, Argentina

Misophonia Labs began as an applied engineering research initiative driven by a critical technological gap: while audio AI has transformed speech recognition and generation, real-time edge processing for sensory and neurological sound sensitivities has remained largely neglected by major consumer hardware manufacturers.

The project focuses on digital signal processing, embedded firmware (C/C++), and deep learning acoustic models, developing bench prototypes to bring clinical-grade relief to everyday social and professional environments.

TA

Tobias Afonso

Founder & Systems Engineer (Engineering Student & Audio ML Researcher)
BahĂ­a Blanca, Buenos Aires Province, Argentina
Questions & Answers

Frequently Asked Questions

Understanding the difference between traditional noise cancelling and selective neural gating.

How is this different from commercial ANC headphones?

Standard ANC operates via analog phase inversion targeted at low-frequency, stationary background hums (engines, HVAC). Misophonia Labs uses digital neural masking to classify and eliminate sudden, non-stationary oral transients (chewing, smacking, clicking) while keeping vocal frequencies unmuted.

Why is latency under 8 milliseconds critical?

If audio passed to your eardrum experiences more than 10-15 ms of delay, you perceive an unnatural comb-filtering echo effect that makes normal conversation disorienting. Our quantized pipeline computes masks in under 3.6 ms, keeping total acoustic latency imperceptible.

Does the device record or send conversations to the cloud?

No. All live audio processing and neural mask inferencing occurs strictly on-device on the embedded micro-controller. No raw audio ever leaves the earphone or streams over the internet.

How does the MCP companion agent interact with the earphone?

The mobile companion app hosts an intelligent conversational agent communicating via Model Context Protocol (MCP). The agent translates user natural language feedback into structured hardware parameters sent over Bluetooth LE to adjust filter floors and frequency cutoffs.

Pilot & Research Access

Join the Early Access Pilot

We are accepting inquiries from clinical audiologists, researchers, and individuals seeking access to upcoming developer prototypes.

Direct contact: contact@misophonialabs.com