IEEE-ACM Transactions on Audio Speech and Language Processing

Papers
(The TQCC of IEEE-ACM Transactions on Audio Speech and Language Processing is 16. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Decorrelation in Feedback Delay Networks367
CET2: Modelling Topic Transitions for Coherent and Engaging Knowledge-Grounded Conversations282
WDEA: The Structure and Semantic Fusion With Wasserstein Distance for Low-Resource Language Entity Alignment255
Towards Generating Diverse Audio Captions via Adversarial Training199
MO-Transformer: Extract High-Level Relationship Between Words for Neural Machine Translation195
DropAttack: A Random Dropped Weight Attack Adversarial Training for Natural Language Understanding181
Reverberant Source Separation Using NTF With Delayed Subsources and Spatial Priors179
Audio-Only Phonetic Segment Classification Using Embeddings Learned From Audio and Ultrasound Tongue Imaging Data142
Generalizing Speaker Verification for Spoof Awareness in the Embedding Space140
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction135
Review of Methods for Automatic Speaker Verification114
Refining Synthesized Speech Using Speaker Information and Phone Masking for Data Augmentation of Speech Recognition94
Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping88
$\mathcal {P}$owMix: A Versatile Regularizer for Multimodal Sentiment Analysis85
Envelope-Based Multichannel Noise Reduction for Cochlear Implant Applications82
Efficient Lightweight Speaker Verification With Broadcasting CNN-Transformer and Knowledge Distillation Training of Self-Attention Maps76
A User-Centric Approach for Deep Residual-Echo Suppression in Double-Talk74
AudioLM: A Language Modeling Approach to Audio Generation66
Learning Discriminative Representations and Decision Boundaries for Open Intent Detection65
The VoxCeleb Speaker Recognition Challenge: A Retrospective65
Enhancing Robustness of Speech Watermarking Using a Transformer-Based Framework Exploiting Acoustic Features64
Attention-Based Speech Enhancement Using Human Quality Perception Modeling61
Representation Learning With Hidden Unit Clustering for Low Resource Speech Applications55
Adaptive Multi-Domain Dialogue State Tracking on Spoken Conversations52
COVID-19 Detection via Fusion of Modulation Spectrum and Linear Prediction Speech Features51
Pronunciation Dictionary-Free Multilingual Speech Synthesis Using Learned Phonetic Representations49
IEEE Signal Processing Society Information49
Emotion Prediction Oriented Method With Multiple Supervisions for Emotion-Cause Pair Extraction48
Complex-Domain Pitch Estimation Algorithm for Narrowband Speech Signals45
ReZero: Region-Customizable Sound Extraction45
Disentangled Text Representation Learning With Information-Theoretic Perspective for Adversarial Robustness45
Spherically Steerable Vector Differential Microphone Arrays43
Implicit Self-Supervised Language Representation for Spoken Language Diarization42
Enhanced Multi-Domain Dialogue State Tracker With Second-Order Slot Interactions41
Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech Processing41
Distinctive and Natural Speaker Anonymization via Singular Value Transformation-Assisted Matrix40
Phrase-Aware Financial Sentiment Analysis Based on Constituent Syntax40
SPEC: Summary Preference Decomposition for Low-Resource Abstractive Summarization40
Predicting Level-Dependent Changes in Concurrent Vowel Scores Using the 2D-CNN Models40
Textless Unit-to-Unit Training for Many-to-Many Multilingual Speech-to-Speech Translation39
Integrated Syntactic and Semantic Tree for Targeted Sentiment Classification Using Dual-Channel Graph Convolutional Network39
End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations38
Source Separation of Piano Concertos Using Musically Motivated Augmentation Techniques38
Knowledge-Guided Transformer for Joint Theme and Emotion Classification of Chinese Classical Poetry37
Label-Correction Capsule Network for Hierarchical Text Classification37
Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach37
Sound Field Estimation Based on Physics-Constrained Kernel Interpolation Adapted to Environment36
Howling Detection and Gain Control for Speech Reinforcement in a Noisy Car Cabin Environment36
Audio-Visual Cross-Attention Network for Robotic Speaker Tracking36
Blind Identification of Ambisonic Reduced Room Impulse Response36
Hate Speech Detection via Dual Contrastive Learning36
Weighted Frequency Smoothing for Enhanced Speaker Localization35
Multi-Layer Combined Frequency and Periodicity Representations for Multi-Pitch Estimation of Multi-Instrument Music35
Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information35
Training a Singing Transcription Model Using Connectionist Temporal Classification Loss and Cross-Entropy Loss35
MusicYOLO: A Vision-Based Framework for Automatic Singing Transcription35
Learning to Perturb for Contrastive Learning of Unsupervised Sentence Representations34
Grouped Feedback Delay Networks With Frequency-Dependent Coupling34
Multi-Grained Evidence Inference for Multi-Choice Reading Comprehension34
FlowHash: Accelerating Audio Search With Balanced Hashing via Normalizing Flow34
TOE: A Grid-Tagging Discontinuous NER Model Enhanced by Embedding Tag/Word Relations and More Fine-Grained Tags33
Tackling Interpretability in Audio Classification Networks With Non-negative Matrix Factorization33
DiCLET-TTS: Diffusion Model Based Cross-Lingual Emotion Transfer for Text-to-Speech — A Study Between English and Mandarin33
Unsupervised Music Source Separation Using Differentiable Parametric Source Models33
Magnitude-Corrected and Time-Aligned Interpolation of Head-Related Transfer Functions33
$F0$ Estimation and Voicing Detection With Cascade Architecture in Noisy Speech32
Neural Multi-Channel and Multi-Microphone Acoustic Echo Cancellation32
Learning Dynamic and Static Representations for Extrapolation-Based Temporal Knowledge Graph Reasoning32
PE-Wav2vec: A Prosody-Enhanced Speech Model for Self-Supervised Prosody Learning in TTS31
Generalized Hyperbolic Tangent Based Random Fourier Conjugate Gradient Filter for Nonlinear Active Noise Control31
Towards Recognition for Radio-Echo Speech in Air Traffic Control: Dataset and a Contrastive Learning Approach31
Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics Retrieval31
Unsupervised Disentanglement Learning Model for Exemplar-Guided Paraphrase Generation31
Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class Information30
Anti-Aliasing Speech DOA Estimation Under Spatial Aliasing Conditions30
Visually Grounded Few-Shot Word Learning in Low-Resource Settings30
Direction of Arrival Estimation of Sound Sources Using Icosahedral CNNs29
Joint Multiscale Cross-Lingual Speaking Style Transfer With Bidirectional Attention Mechanism for Automatic Dubbing29
Decomposition-Based Wiener Filter Using the Kronecker Product and Conjugate Gradient Method29
Scalable-Complexity Steered Response Power Based on Low-Rank and Sparse Interpolation29
CoNeTTE: An Efficient Audio Captioning System Leveraging Multiple Datasets With Task Embedding28
Configurable EBEN: Extreme Bandwidth Extension Network to Enhance Body-Conducted Speech Capture28
DiaPer: End-to-End Neural Diarization With Perceiver-Based Attractors28
Attention-Based Encoder-Decoder End-to-End Neural Diarization With Embedding Enhancer28
Unified Instance and Knowledge Alignment Pretraining for Aspect-Based Sentiment Analysis28
Multi-Channel Speech Separation Using Spatially Selective Deep Non-Linear Filters27
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting27
Query-Efficient Black-Box Adversarial Attacks on Automatic Speech Recognition27
Enhancing Conformer-Based Sound Event Detection Using Frequency Dynamic Convolutions and BEATs Audio Embeddings27
Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech Segmentation27
A New Diffusion Filtered-X Affine Projection Algorithm: Performance Analysis and Application in Windy Environment26
ASiT: Local-Global Audio Spectrogram Vision Transformer for Event Classification26
Audio-Visual End-to-End Multi-Channel Speech Separation, Dereverberation and Recognition26
An Analysis of Traditional Noise Power Spectral Density Estimators Based on the Gaussian Stochastic Volatility Model26
Refining History for Future-Aware Neural Machine Translation26
Active Discovering New Slots for Task-Oriented Conversation25
The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance25
A Perceptually Evaluated Signal Model: Collisions Between a Vibrating Object and an Obstacle25
Distance Metric-Based Open-Set Domain Adaptation for Speaker Verification25
TDFNet: Transformer-Based Deep-Scale Fusion Network for Multimodal Emotion Recognition25
The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation25
Optimal Modal Decomposition for Directionally Biased Sound Field Recording25
U-Shaped Transformer With Frequency-Band Aware Attention for Speech Enhancement25
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization25
TriSAT: Trimodal Representation Learning for Multimodal Sentiment Analysis24
Contrastive Learning for Target Speaker Extraction With Attention-Based Fusion24
Higher-Order Stereophony24
Minimum Processing Near-End Listening Enhancement24
BEHM-GAN: Bandwidth Extension of Historical Music Using Generative Adversarial Networks24
Enhancing Paraphrase Question Generation With Prior Knowledge24
EchoScan: Scanning Complex Room Geometries via Acoustic Echoes23
A CTC Alignment-Based Non-Autoregressive Transformer for End-to-End Automatic Speech Recognition23
Meta-AF: Meta-Learning for Adaptive Filters23
Latent-Domain Predictive Neural Speech Coding23
Coefficients-Switched Normalized Least-Mean- Squares Adaption in Echo Canceler of Sparse-Echo-Path23
Towards Improved Objective Perceptual Audio Quality Assessment - Part 1: A Novel Data-Driven Cognitive Model23
Consonant-Vowel Transition Models Based on Deep Learning for Objective Evaluation of Articulation23
Hybrid-Frequency-Resolution Adaptive Kalman Filter for Online Identification of Long Acoustic Responses With Low Input-Output Latency23
Adversarial Multi-Task Learning for Mandarin Prosodic Boundary Prediction With Multi-Modal Embeddings23
MRC-PASCL: A Few-Shot Machine Reading Comprehension Approach via Post-Training and Answer Span-Oriented Contrastive Learning23
Heterogeneous-Graph Reasoning With Context Paraphrase for Commonsense Question Answering22
ZMM-TTS: Zero-Shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-Supervised Discrete Speech Representations22
A Flexible Architecture Using Temporal, Spatial and Semantic Correlation-Based Algorithms for Story Segmentation of Broadcast News22
On the Predictive Power of Objective Intelligibility Metrics for the Subjective Performance of Deep Complex Convolutional Recurrent Speech Enhancement Networks22
A New Virtual Tracking Sub-Algorithm Based Hybrid Active Control System for Narrowband Noise With Impulsive Interference22
MOSA: Music Motion With Semantic Annotation Dataset for Cross-Modal Music Processing22
List of Reviewers21
Improving Non-Autoregressive Translation Quality With Pretrained Language Model, Embedding Distillation and Upsampling Strategy for CTC21
Multi-Source Localization Using Optimized Time-Frequency Representation and Sparsity Component Analysis21
Statistical Analysis for Speaker Recognition Evaluation With Data Dependence and Three Score Distributions21
Multi-Source Discriminant Subspace Alignment for Cross-Domain Speech Emotion Recognition21
SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual Data21
Speech Enhancement and Dereverberation With Diffusion-Based Generative Models21
Empathetic Response Generation Based on Plug-and-Play Mechanism With Empathy Perturbation21
A Composite T60 Regression and Classification Approach for Speech Dereverberation21
Block-Based Perceptually Adaptive Sound Zones With Reproduction Error Constraints20
Dynamic Prompt-Driven Zero-Shot Relation Extraction20
Cross-Domain Aspect-Based Sentiment Classification With Tripartite Graph Modeling20
Decomposed Meta-Learning for Few-Shot Sequence Labeling20
En-HACN: Enhancing Hybrid Architecture With Fast Attention and Capsule Network for End-to-end Speech Recognition20
Operation-Augmented Numerical Reasoning for Question Answering20
RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion Recognition20
Joint Maximum Likelihood Estimation of Microphone Array Parameters for a Reverberant Single Source Scenario19
Interrelate Training and Clustering for Online Speaker Diarization19
FTDKD: Frequency-Time Domain Knowledge Distillation for Low-Quality Compressed Audio Deepfake Detection19
Gradformer: A Framework for Multi-Aspect Multi-Granularity Pronunciation Assessment19
Segment-Less Continuous Speech Separation of Meetings: Training and Evaluation Criteria19
Uncertainty-Driven Knowledge Distillation for Language Model Compression19
Multi-Cue Guided Semi-Supervised Learning Toward Target Speaker Separation in Real Environments19
LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification19
Text-Inductive Graphone-Based Language Adaptation for Low-Resource Speech Synthesis19
EmoInt-Trans: A Multimodal Transformer for Identifying Emotions and Intents in Social Conversations19
JMS-QA: A Joint Hierarchical Architecture for Mental Health Question Answering19
FxLMS/F Based Tap Decomposed Adaptive Filter for Decentralized Active Noise Control System19
Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models19
Joint Dual Learning With Mutual Information Maximization for Natural Language Understanding and Generation in Dialogues19
Towards Unified Multi-Domain Machine Translation With Mixture of Domain Experts18
Decoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent Recognition18
Artist Similarity Based on Heterogeneous Graph Neural Networks18
Selective Acoustic Feature Enhancement for Speech Emotion Recognition With Noisy Speech18
NoiseBandNet: Controllable Time-Varying Neural Synthesis of Sound Effects Using Filterbanks18
Enhancing Low-Resource NLP by Consistency Training With Data and Model Perturbations18
JoinER-BART: Joint Entity and Relation Extraction With Constrained Decoding, Representation Reuse and Fusion18
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR18
Multi-Task Attentive Residual Networks for Argument Mining17
Automatic Detection of Speech Sound Disorder in Cantonese-Speaking Pre-School Children17
Direct and Residual Subspace Decomposition of Spatial Room Impulse Responses17
Data-Centric Methods for Environmental Sound Classification With Limited Labels17
Controllable Dialogue Generation With Disentangled Multi-Grained Style Specification and Attribute Consistency Reward17
Distributed Microphone Array Localization Problem via SDP-SOCP Method17
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation17
Adaptive Ensemble Self-Distillation With Consistent Gradients for Fast Inference of Pretrained Language Models16
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks16
STFF-SM: Steganalysis Model Based on Spatial and Temporal Feature Fusion for Speech Streams16
WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation Learning16
Assessing the Generalization Gap of Learning-Based Speech Enhancement Systems in Noisy and Reverberant Environments16
A Two-Stage Deep Representation Learning-Based Speech Enhancement Method Using Variational Autoencoder and Adversarial Training16
Music Source Separation With Band-Split RNN16
SinTechSVS: A Singing Technique Controllable Singing Voice Synthesis System16
Recent Trends in Deep Learning Based Textual Emotion Cause Extraction16
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering16
Statistically Guided Near-End Speech Intelligibility Improvement Through Voice Transformation and Transfer Learning16
A Multi-Level Supervised Contrastive Learning Framework for Low-Resource Natural Language Inference16
Interpretable Spectrum Transformation Attacks to Speaker Recognition Systems16
Boosting Cross-Domain Speech Recognition With Self-Supervision16
Localization-Driven Speech Enhancement in Noisy Multi-Speaker Hospital Environments Using Deep Learning and Meta Learning16
Multi-Channel Conversational Speaker Separation via Neural Diarization16
0.50894594192505