← All workCH.02 · 2025

Speech Enhancement for Defense

Real-time speech denoising small enough to run on the edge.

PyTorchConv-TasNetONNXDSP

A speech enhancement model that pulls intelligible speech out of heavy, unpredictable noise, built to run in real time on constrained edge hardware. A lightweight Conv-TasNet is trained on dynamically mixed data and shipped as a quantized ONNX graph.

Approach

  1. Lightweight Conv-TasNet: a time-domain separation network trimmed for latency and memory.
  2. Dynamic-SNR data mixing: every batch remixes clean speech and noise at fresh signal-to-noise ratios, so the model never memorises one noise floor.
  3. Augmentation for field conditions: room reverb and clipping, to mimic real radios and microphones.
  4. Trained directly on SI-SNR loss; evaluated on SNR, STOI (intelligibility) and PESQ (perceived quality).
  5. An optional LMS adaptive filter stage for stationary interference.
  6. Exported to ONNX with int8 quantization for real-time inference on the edge.