AnomSepBridging Optimization-Based Separation and Deep Neural Refinement for Anomaly Detection with Listenable Explanations

Koki ShodaJun Younes Louhi KasaharaQi AnAtsushi Yamashita

The University of Tokyo

The idea

Hear the sound behind an anomaly score.

An anomaly score flags a departure from everyday sounds. AnomSep also produces an audio waveform that lets a listener inspect the evidence behind that score. An anomalous sound deviates from an environment’s normal sound distribution; an anomaly does not necessarily indicate danger.

01

Separate

Nonnegative Novelty Extraction (NNE) estimates normal and anomalous components using a dictionary learned from normal recordings.

02

Refine

Flow matching in a pretrained audio latent space refines the initial estimates. Normal Region Exclusion selects surrogate anomalous audio for adaptation.

03

Detect and listen

The energy of the refined anomalous component provides an anomaly score. The same component can be played back as a listenable explanation.

AnomSep pipeline: a mixture enters NNE; the preliminary normal-sound and anomalous-sound estimates are encoded, refined by flow matching conditioned on the mixture, and decoded into refined estimates.
AnomSep initializes neural refinement with the source estimates produced by NNE.

Listening room

From a mixture to listenable explanations

Compare four methods on five mixtures containing anomalous sounds and five normal-only recordings. All method outputs below are estimates of the anomalous component.

Start with the mixture and references. Then compare the estimated anomalous sounds. Successful separation preserves the anomalous event and suppresses the normal background. A normal-only example should produce a quiet anomalous-sound estimate.

CLAP is audio embedding cosine similarity to the anomalous reference; SAJ is the automated SAM Audio Judge overall score. Higher is better for both. Scores are undefined for silent anomalous references in normal-only examples and are shown as N/A.

Loading listening examples…

Consistent playback. No track is normalized independently. A shared attenuation, when needed, keeps every track within an example below full scale and preserves relative amplitudes. Keep the player volume unchanged when comparing methods.

−80 to 0.0 dB · one spectrogram scale per example · 0.0–8.0 kHz

Spectrogram levels are relative to the maximum magnitude across all reference and estimated tracks in the selected example. Detection labels are the saved predictions from each method’s tuning-selected threshold.

Use AnomSep

Code and pretrained models

The release provides reusable AnomSep and Normal Region Exclusion classes, scene-specific model weights, and reproducible demo selection.

Original source code is available for noncommercial research only. See the LICENSE.