IEEE Transactions on Multimedia

Papers
(The TQCC of IEEE Transactions on Multimedia is 16. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Improving Vision Anomaly Detection With the Guidance of Language Modality1006
Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo Selection547
Rethinking Video Sentence Grounding From a Tracking Perspective With Memory Network and Masked Attention398
Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective366
FoodSAM: Any Food Segmentation261
Online Low-Light Sand-Dust Video Enhancement Using Adaptive Dynamic Brightness Correction and a Rolling Guidance Filter228
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization207
SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud Analysis204
ViDR-GNN: Vision Implicit Discriminative Reorganization Graph Neural Networks202
Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding196
Few-Shot Generative Model Adaptation via Style-Guided Prompt195
Weakly-Supervised Video Object Grounding via Learning Uni-Modal Associations191
Towards Substation Semantic Segmentation: A benchmark dataset and a cross-attention embedded hierarchical network184
HRVFusion: Video-based Long-Term Heart Rate Variability Measurement with Conditional Diffusion Models168
Revisiting the Adversarial Transferability: Towards a Perspective of Semantic Preservation162
Exploring Kernel Transformations for Implicit Neural Representations160
Posture-Movement-Frequency-Enhanced Graph Convolutional Network for Gait Emotion Recognition158
Mask-Aware Kernel Learning for Action Recognition156
LMAgent: A Large-scale Multimodal Agents Society for Multi-user Simulation156
Bias-Correction Feature Learner for Semi-Supervised Instance Segmentation155
Mix-Based Training Strategies for Learning Implicit Neural Representations154
Bidirectional Translation Between UHD-HDR and HD-SDR Videos153
Optimal Transport-Based Patch Matching for Image Style Transfer151
Neighborhood Contrastive Transformer for Change Captioning151
Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning144
Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait Recognition142
PropMambaSR: Lightweight Image Super-Resolution with Propagation State Space Model139
Adaptive Weight Generator for Multi-Task Image Recognition by Task Grouping Prompt138
Semantic-Aware Triplet Loss for Image Classification136
Semantic Dual-Adversarial Network for Blended-Target Domain Adaptation136
DWSF-Net: A Dynamic Wavelet-based Spatial-frequency Fusion Network for Multispectral Object Detection135
AMS-Net: Adaptive Multi-Scale Network for Image Compressive Sensing135
Late Fusion Multiple Kernel Clustering With Local Kernel Alignment Maximization134
Disaggregation Distillation for Person Search132
Rényi Entropy Induced Efficient and Balanced One-Step Multi-View Clustering132
Vision-Controllable Language Model for Image-Guided Story Ending Generation128
Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics Assessment125
Semi-Supervised Contrastive Learning With Similarity Co-Calibration123
Distributed Deep Point Cloud Feature Compression for Vehicle-to-Vehicle Cooperative Perception120
Guided Image-to-Image Translation by Discriminator-Generator Communication119
Weakly-Supervised 3D Visual Grounding Based on Visual Language Alignment119
One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing118
MHRN: A Multimodal Hierarchical Reasoning Network for Topic Detection115
SCSP: An Unsupervised Image-to-Image Translation Network Based on Semantic Cooperative Shape Perception115
BMB: Balanced Memory Bank for Long-Tailed Semi-Supervised Learning114
Unsupervised Learning-Based Framework for Deepfake Video Detection114
Long Video Understanding With Learnable Retrieval in Video-Language Models113
Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames112
Transferable Backdoor Attack on Any CLIP Model With Any Target Class by Pre-Trained Hack Network111
Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude Similarity111
Asymptotics-Aware Multi-View Subspace Clustering110
Structured Graph Reasoning for Traffic Anomaly Detection108
Self-Guided Discriminative Locality Preserving Projections107
Vulnerability of Feature Extractors in 2D Image-Based 3D Object Retrieval105
Dynamic Mosaics: Saliency-Guided Adaptive Masking for Occluded Person Re-Identification104
MGKsite: Multi-Modal Knowledge-Driven Site Selection via Intra and Inter-Modal Graph Fusion104
Beyond Simple Extraction: Unleashing the Potential of Encoder Interaction in Few-Shot Segmentation103
Interpretable Graph Convolutional Network for Multi-View Semi-Supervised Learning102
TFBF: Temporal-Frequency Bidirectional Fusion for Action Quality Assessment100
Outliers Adaptation Exploration and Centroids Matching Label Refinement for Unsupervised Person Re-Identification100
GLCT: A Novel Global-Local Constraint for Unpaired Image-to-Image Translation100
SkyML: A MLaaS Federation Design for Multicloud-Based Multimedia Analytics98
Cross-modal Semantic Relevance is An Efficient Gatekeeper for Audio-Visual Video Parsing98
ICE: Interactive 3D Game Character Facial Editing via Dialogue97
Ensemble Prototype Networks for Unsupervised Cross-Modal Hashing With Cross-Task Consistency96
Siamese Alignment Network for Weakly Supervised Video Moment Retrieval96
MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition96
Anomaly-Led Prompting Learning Caption Generating Model and Benchmark96
Distilling Multi-View Diffusion Models Into 3D Generators94
Skeleton-Based Action Recognition With Select-Assemble-Normalize Graph Convolutional Networks94
Adversarial 3D-to-Real Watermarking: Revealing Invisible Messages Hidden in Complexly Distorted Surfaces93
BASNet: Boundary Assisted Network for Image Splicing Forgery Detection91
Scale Up Composed Image Retrieval Learning via Modification Text Generation91
Pixel Bleach Network for Detecting Face Forgery Under Compression90
Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image Segmentation90
3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature Enhancement88
$\rm {M}^{2}\rm {C-EvDet}$: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection88
Progressive Local Filter Pruning for Image Retrieval Acceleration87
XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework86
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation With Interpretability85
Semi-Supervised Domain Adaptation for Major Depressive Disorder Detection84
Dynamic Contrastive Distillation for Image-Text Retrieval83
Feature First: Advancing Image-Text Retrieval Through Improved Visual Features83
Deep Semantic-Consistent Penalizing Hashing for Cross-Modal Retrieval82
Improving Pre-Trained Model-Based Speech Emotion Recognition From a Low-Level Speech Feature Perspective82
Semi-Supervised Domain Adaptation via Joint Transductive and Inductive Subspace Learning82
Perceptual Image Hashing Using Feature Fusion of Orthogonal Moments80
CVKD-UDA: Cross-View Knowledge Distillation for 3D Unsupervised Domain Adaptive Segmentation79
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition79
Hierarchical Equalization Loss for Long-Tailed Instance Segmentation78
SLCGC: A lightweight Self-supervised Low-Pass Contrastive Graph Clustering Network for Hyperspectral Images77
Adaptive HEVC Video Steganography With High Performance Based on Attention-Net and PU Partition Modes77
PhotoHelper: Portrait Photographing Guidance Via Deep Feature Retrieval and Fusion77
Towards Temporal Event Detection: A Dataset, Benchmarks and Challenges74
Towards Neural Codec-Empowered 360$^\circ$ Video Streaming: A Saliency-Aided Synergistic Approach74
Image-Based Structured Vehicle Behavior Analysis Inspired by Interactive Cognition73
Cps-STS: Bridging the Gap Between Content and Position for Coarse-Point-Supervised Scene Text Spotter72
ALCER3D: Adaptive Learning Constraints for Enhanced Retrieval of Complex Indoor 3D Scenarios72
Supervised Contrastive Learning for Indoor Point Cloud Oversegmentation71
DEHand: Deformable Encoding for Photo-Realistic Free-View and Free-Pose Hand Rendering70
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios70
Cross-Domain Sample Relationship Learning for Facial Expression Recognition68
Semi-Supervised Authentically Distorted Image Quality Assessment With Consistency-Preserving Dual-Branch Convolutional Neural Network68
JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging Perspective68
Reconstructed Graph Constrained Auto-Encoders for Multi-View Representation Learning68
Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models67
Motion Direction Awareness: A Biomimetic Dynamic Capture Mechanism for Video Prediction67
Depth Map Super-Resolution via Deep Cross-Modality and Cross-Scale Guidance67
High Specificity Guided Cross-Domain Few-Shot Segmentation66
Vulnerabilities in AI-Generated Image Detection: The Challenge of Adversarial Attacks66
RUL: Region Uncertainty Learning for Robust Face Recognition66
Rate-Adaptive Neural Network for Image Compressive Sensing66
Improving Out-of-Distribution Generalization on Point Clouds with Cross-Domain Adversarial Distillation65
Boosting Universal Adversarial Attack on Deep Neural Networks65
Enhanced Context Mining and Filtering for Learned Video Compression65
Investigating the Effective Dynamic Information of Spectral Shapes for Audio Classification65
Reliable Multi-View Clustering with Graph Neural Network64
Video Instance Segmentation by Instance Flow Assembly64
Saliency-Aware Adversarial Attacks on Visual Trackers63
Human-Centric Behavior Description in Videos: New Benchmark and Model63
Cooperative Bargaining Game Based Adaptive Video Multicast Over Mobile Edge Networks63
Denoised Semantic Features for Local Consistent No-Reference Image Quality Assessment62
Personalized Fashion Recommendation With Discrete Content-Based Tensor Factorization62
EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation62
A Multidimensional Media Adaptation Framework for Live Holographic Communication61
Multimodal Progressive Modulation Network for Micro-Video Multi-Label Classification61
Spatial-Temporal Saliency Guided Unbiased Contrastive Learning for Video Scene Graph Generation61
CMANet: Context-Aware Mutual Attention Network for Referring Image Segmentation60
SDE2D: Semantic-Guided Discriminability Enhancement Feature Detector and Descriptor60
Exploring Kernel-Based Texture Transfer for Pose-Guided Person Image Generation60
Enhancing Representation Inversion and Alignment for Zero-Shot Composed Image Retrieval60
Can Machines Generate Personalized Music? A Hybrid Favorite-Aware Method for User Preference Music Transfer60
MMIFN: A Multi-Modal Interactive Fusion Network for Omnidirectional Image Quality Assessment60
DREAMT: Diversity Enlarged Mutual Teaching for Unsupervised Domain Adaptive Person Re-Identification59
Wavelet-Domain Masked Image Modeling for Color-Consistent HDR Video Reconstruction59
Look&listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement58
Compositional Text-to-Image Synthesis With Training-Free Layout-Guided Diffusion58
Velocity First? Rethinking 3D Object Detection with 4D Millimeter Wave Radar58
Show, Tell and Rephrase: Diverse Video Captioning via Two-Stage Progressive Training58
Dynamic Strategy Prompt Reasoning for Emotional Support Conversation58
Exploring Basic Expression Representation for Compound Facial Expression Recognition58
Prototypical Bidirectional Adaptation and Learning for Cross-Domain Semantic Segmentation57
Generalizing Beyond Patterns: Dynamic Moment Query Recalibrating for Out-of-Distribution Video Temporal Localization57
Bidirectional Prototype-Reward Co-Evolution for Test-Time Adaptation of Vision-Language Models57
Exploring Local and Global Consistent Correlation on Hypergraph for Rotation Invariant Point Cloud Analysis56
TPE-ADE: Thumbnail-Preserving Encryption Based on Adaptive Deviation Embedding for JPEG Images56
VOLTER: Visual Collaboration and Dual-Stream Fusion for Scene Text Recognition55
Motion Deblur by Learning Residual From Events55
HP-C4D: A Fast Camera and 4D Radar Fusion Framework With Height Prediction for 3D Object Detection55
UniCrossGait: Unified Cross-Modal Gait Recognition Based on Knowledge Distillation55
A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification54
RetinexGS: Enhancing 3D Gaussian Splatting for Low-Light Scenes Via Retinex-Guided Decomposition54
Deep Unfolding Network for Image Compressed Sensing by Content-Adaptive Gradient Updating and Deformation-Invariant Non-Local Modeling54
STNet: Scale Tree Network With Multi-Level Auxiliator for Crowd Counting54
HSV-Driven Illumination-Aware Iterative Network for Unsupervised Low-Light Enhancement54
FedSH: Towards Privacy-Preserving Text-Based Person Re-Identification54
CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event Cameras54
Sparse Transformer for Ultra-Sparse Sampled Video Compressive Sensing54
Universal Infrared Image Nonuniformity Correction via Stripe-Aware Attention Network54
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding54
Ocean's Duality: Physics-Driven Dual-Branch Framework for Underwater Vision Enhancement53
Twin Tensor Learning for Consistency and Inconsistency: A Unified Affinity Learning Framework for Multi-View Clustering53
RSNet: Relation Separation Network for Few-Shot Similar Class Recognition53
Sounding Depressed? Personalized Deep Learning Model for Depression Detection From Speech and Text53
Primary Code Guided Targeted Attack against Cross-modal Hashing Retrieval53
Sentiment-Enhanced Graph-Based Sarcasm Explanation in Dialogue53
No-Reference Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded Point Clouds52
Video-to-Music Recommendation Using Temporal Alignment of Segments52
Action-Responsive Contrastive Network for Fine-Grained Skeleton-Based Action Recognition52
RA-SCIC: Region-Aware Screen Content Image Coding Towards Fidelity and Efficiency52
Multi-View User Preference Modeling for Personalized Text-to-Image Generation52
Pedestrian Trajectory Prediction Based on Social Interactions Learning With Random Weights52
IEIRNet: Inconsistency Exploiting Based Identity Rectification for Face Forgery Detection51
Probabilistic Temporal Masked Attention for Cross-View Online Action Detection51
Style-Agnostic Representation Learning for Visible-Infrared Person Re-Identification51
Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object Detection51
Test-Time Model Adaptation for Visual Question Answering With Debiased Self-Supervisions51
FFFN: Frame-By-Frame Feedback Fusion Network for Video Super-Resolution50
Dual Representation Aggregation Network for Blind Image Super-Resolution via Iterative Bi-level Optimization50
C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image Compression50
Knowledge Distillation-Based Domain-Invariant Representation Learning for Domain Generalization50
REDEditing: Relationship-Driven Precise Backdoor Poisoning on Text-to-Image Diffusion Models50
High Fidelity Face-Swapping With Style ConvTransformer and Latent Space Selection50
SegTrans: Transferable Adversarial Examples for Segmentation Models50
Anchor-Guided Discrete Multi-View Clustering50
Exploiting EfficientSAM and Temporal Coherence for Audio-Visual Segmentation49
Underwater Image Enhancement With Cascaded Contrastive Learning49
DA-Net: Density-Aware 3D Object Detection Network for Point Clouds49
Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image Enhancement49
RD-VTA: Rule-Data Guided Video-to-Audio Generation for Fine-Grained Footstep Sound49
Improving Fine-Grained Image Classification With Multimodal Information49
GLFF: Global and Local Feature Fusion for AI-Synthesized Image Detection49
Multimodal Sentiment Analysis With Image-Text Interaction Network49
Low-Light Image Enhancement via Self-Reinforced Retinex Projection Model49
Towards Region-Aware Finer Self-Supervised Learning for Fine-Grained Visual Recognition48
FGDNet: Fine-Grained Detection Network Towards Face Anti-Spoofing48
Reordered $k$-Means: A New Baseline for View-Unaligned Multi-View Clustering48
MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures48
Augment One With Others: Generalizing to Unforeseen Variations for Visual Tracking48
OpenSlot: Mixed Open-Set Recognition With Object-Centric Learning48
Bridging the Short-Term and Long-Term Gap: A Cross-Task Continuous Learning Person Re-Identification Problem47
Multi-View Depth Estimation With Uncertainty Constraints for Virtual-Real Occlusion47
Semantics Alternating Enhancement and Bidirectional Aggregation for Referring Video Object Segmentation47
Rethinking the Role of Vector Quantization for Blind Image Restoration47
FOF-X: Towards Real-time Detailed Human Reconstruction from a Single Image47
DDGA: Domain Distance Guided Active Domain Adaptation for Unpaired Super-Resolution47
Dense Video Captioning With Early Linguistic Information Fusion47
Blind Video Quality Assessment at the Edge47
Unleash the Power of Vision-Language Models by Visual Attention Prompt and Multimodal Interaction47
Flow Guidance Deformable Compensation Network for Video Frame Interpolation47
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection47
CenterTube: Tracking Multiple 3D Objects With 4D Tubelets in Dynamic Point Clouds47
Underwater Adaptive Video Transmissions Using MIMO-Based Software-Defined Acoustic Modems46
SSPNet: Predicting Visual Saliency Shifts46
Edge-Assisted Massive Video Delivery Over Cell-Free Massive MIMO46
Towards a Multi-Granulated Statistical Framework for Human–Machine Collaboration in Image Classification46
Tensorformer: Normalized Matrix Attention Transformer for High-Quality Point Cloud Reconstruction46
Graph Convolutional Network With Unknown Class Number46
RaFPN: Relation-Aware Feature Pyramid Network for Dense Image Prediction46
Unsupervised Deepfake Detection via Camera Source Clustering and Temporal-Spatial Features46
Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment46
VRTNet: Vector Rectifier Transformer for Two-View Correspondence Learning46
Compression of Plenoptic Point Cloud Attributes Using 6-D Point Clouds and 6-D Transforms46
SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual Grounding46
Category-Contrastive Fine-Grained Crowd Counting and Beyond46
Cross-Modality Feature Fusion for Forward-Looking Sonar Image Segmentation in Complex Underwater Environments46
Simultaneously Training and Compressing Vision-and-Language Pre-Training Model46
Inexactly Matched Referring Expression Comprehension With Rationale46
Scene Graph Knowledge Enhanced Hashing with Contrastive Learning for Image-Text Retrieval46
Soundscape Captioning Using Sound Affective Quality Network and Large Language Model46
Progressive Learning of Instance-Level Proxy Semantics for Few-Shot Action Recognition45
CMI-Net: Cross-View Message Token Interaction Network for 3D Shape Recognition45
Progressive Learning Model for Big Data Analysis Using Subnetwork and Moore-Penrose Inverse45
Point Cloud Soft Multicast for Untethered XR Users45
Visibility-Based Geometry Pruning of Neural Plenoptic Scene Representations45
Instruction-Driven 3D Facial Expression Generation and Transition45
Noise Aware Audio-Visual Speech Denoising45
MVL-Net: Pairwise Learning for Multi-View Multiple People Labelling45
Tuning-Free High-Resolution Video Diffusion With Spatial-Temporal Latent Grouping45
Interpretable Multi-View Representation Learning Towards Complex Scenes: From Homogeneity to Heterogeneity45
Question Understanding and Temporality Guiding for Video Question Answering45
MPPM: A Mobile-Efficient Part Model for Object re-ID44
Low-Light Image Enhancement With SAM-Based Structure Priors and Guidance44
Face De-Occlusion With Deep Cascade Guidance Learning44
CNIE: Content-Aware Non-Transferable Information Extraction for Fine-Grained Visual Categorization44
Neural-Enhanced Rate Adaptation and Computation Distribution for Emerging mmWave Multi-User 3D Video Streaming Systems44
CLCT: Complementary Local Consensus Transformer for Two-View Correspondence Pruning44
Bidirectional Maximum Entropy Training With Word Co-Occurrence for Video Captioning44
Indistinguishability Analysis of JPEG Image Encryption Schemes44
Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image Decomposition44
0.87858200073242