NovelSepBridging Optimization-Based Separation and Deep Neural Refinement for Novelty Detection with Listenable Explanations

Koki ShodaJun Younes Louhi KasaharaQi AnAtsushi Yamashita

The University of Tokyo

The idea

Hear what made a sound novel.

A novelty score flags a departure from everyday sounds. NovelSep also produces an audio waveform that lets a listener inspect the evidence behind that score. A novel sound is new relative to an environment’s normal sound distribution; novelty does not necessarily indicate danger.

01

Separate

Nonnegative Novelty Extraction (NNE) estimates normal and novel components using a dictionary learned from normal recordings.

02

Refine

Flow matching in a pretrained audio latent space refines the initial estimates. Normal Region Exclusion selects surrogate novel audio for adaptation.

03

Detect and listen

The energy of the refined novel component provides a novelty score. The same component can be played back as a listenable explanation.

NovelSep pipeline: a mixture enters NNE; the preliminary normal and novel estimates are encoded, refined by flow matching conditioned on the mixture, and decoded into refined estimates.
NovelSep initializes neural refinement with the source estimates produced by NNE.

Listening room

From a mixture to listenable explanations

Compare four methods on five novel-containing mixtures and five normal-only recordings. All method outputs below are estimates of the novel component.

Start with the mixture and references. Then compare the estimated novel sounds. Successful separation preserves the novel event and suppresses the normal background. A normal-only example should produce a quiet novel estimate.

CLAP is audio embedding cosine similarity to the novel reference; SAJ is the automated SAM Audio Judge overall score. Higher is better for both. Scores are undefined for silent normal-only novel references and are shown as N/A.

Loading listening examples…

Consistent playback. No track is normalized independently. A shared attenuation, when needed, keeps every track within an example below full scale and preserves relative amplitudes. Keep the player volume unchanged when comparing methods.

−80 to 0.0 dB · one spectrogram scale per example · 0.0–8.0 kHz

Spectrogram levels are relative to the maximum magnitude across all reference and estimated tracks in the selected example. Detection labels are the saved predictions from each method’s tuning-selected threshold.

Use NovelSep

Code and pretrained models

The release provides reusable NovelSep and Normal Region Exclusion classes, scene-specific model weights, and reproducible demo selection.

Original source code is available for noncommercial research only. See the LICENSE.