Comparing neural network architectures for non-intrusive speech quality prediction

Non-intrusive speech quality predictors evaluate speech quality without the use of a reference signal, making them useful in many practical applications. Recently, neural networks have shown the best performance for this task. Two such models in the literature are the convolutional neural network ba...

Full description

Saved in:
Bibliographic Details
Published in:Speech communication Vol. 165; p. 103123
Main Authors: Schill, Leif Førland, Piechowiak, Tobias, Laroche, Clément, Mowlaee, Pejman
Format: Journal Article
Language:English
Published: Elsevier B.V 01-11-2024
Subjects:
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:Non-intrusive speech quality predictors evaluate speech quality without the use of a reference signal, making them useful in many practical applications. Recently, neural networks have shown the best performance for this task. Two such models in the literature are the convolutional neural network based DNSMOS and the bi-directional long short-term memory based Quality-Net, which were originally trained to predict subjective targets and intrusive PESQ scores, respectively. In this paper, these two architectures are trained on a single dataset, and used to predict the intrusive ViSQOL score. The evaluation is done on a number of test sets with a variety of mismatch conditions, including unseen speech and noise corpora, and common voice over IP distortions. The experiments show that the models achieve similar predictive ability on the training distribution, and overall good generalization to new noise and speech corpora. Unseen distortions are identified as an area where both models generalize poorly, especially DNSMOS. Our results also suggest that a pervasiveness of ambient noise in the training set can cause problems when generalizing to certain types of noise. Finally, we detail how the ViSQOL score can have undesirable dependencies on the reference pressure level and the voice activity level. •Non-intrusive speech quality prediction of VISQOL target using neural networks.•CNN and BLSTM architectures compared on one dataset.•Good generalization to unseen speech and noise corpora.•Poor generalization to unseen distortions.•The VISQOL score is dependent on reference level and voice activity level.
ISSN:0167-6393
DOI:10.1016/j.specom.2024.103123