IEEE Transactions on Multimedia

Papers
(The H4-Index of IEEE Transactions on Multimedia is 83. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Improving Vision Anomaly Detection With the Guidance of Language Modality1006
Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo Selection547
Rethinking Video Sentence Grounding From a Tracking Perspective With Memory Network and Masked Attention398
Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective366
FoodSAM: Any Food Segmentation261
Online Low-Light Sand-Dust Video Enhancement Using Adaptive Dynamic Brightness Correction and a Rolling Guidance Filter228
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization207
SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud Analysis204
ViDR-GNN: Vision Implicit Discriminative Reorganization Graph Neural Networks202
Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding196
Few-Shot Generative Model Adaptation via Style-Guided Prompt195
Weakly-Supervised Video Object Grounding via Learning Uni-Modal Associations191
Towards Substation Semantic Segmentation: A benchmark dataset and a cross-attention embedded hierarchical network184
HRVFusion: Video-based Long-Term Heart Rate Variability Measurement with Conditional Diffusion Models168
Revisiting the Adversarial Transferability: Towards a Perspective of Semantic Preservation162
Exploring Kernel Transformations for Implicit Neural Representations160
Posture-Movement-Frequency-Enhanced Graph Convolutional Network for Gait Emotion Recognition158
LMAgent: A Large-scale Multimodal Agents Society for Multi-user Simulation156
Mask-Aware Kernel Learning for Action Recognition156
Bias-Correction Feature Learner for Semi-Supervised Instance Segmentation155
Mix-Based Training Strategies for Learning Implicit Neural Representations154
Bidirectional Translation Between UHD-HDR and HD-SDR Videos153
Optimal Transport-Based Patch Matching for Image Style Transfer151
Neighborhood Contrastive Transformer for Change Captioning151
Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning144
Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait Recognition142
PropMambaSR: Lightweight Image Super-Resolution with Propagation State Space Model139
Adaptive Weight Generator for Multi-Task Image Recognition by Task Grouping Prompt138
Semantic-Aware Triplet Loss for Image Classification136
Semantic Dual-Adversarial Network for Blended-Target Domain Adaptation136
AMS-Net: Adaptive Multi-Scale Network for Image Compressive Sensing135
DWSF-Net: A Dynamic Wavelet-based Spatial-frequency Fusion Network for Multispectral Object Detection135
Late Fusion Multiple Kernel Clustering With Local Kernel Alignment Maximization134
Disaggregation Distillation for Person Search132
Rényi Entropy Induced Efficient and Balanced One-Step Multi-View Clustering132
Vision-Controllable Language Model for Image-Guided Story Ending Generation128
Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics Assessment125
Semi-Supervised Contrastive Learning With Similarity Co-Calibration123
Distributed Deep Point Cloud Feature Compression for Vehicle-to-Vehicle Cooperative Perception120
Guided Image-to-Image Translation by Discriminator-Generator Communication119
Weakly-Supervised 3D Visual Grounding Based on Visual Language Alignment119
One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing118
MHRN: A Multimodal Hierarchical Reasoning Network for Topic Detection115
SCSP: An Unsupervised Image-to-Image Translation Network Based on Semantic Cooperative Shape Perception115
BMB: Balanced Memory Bank for Long-Tailed Semi-Supervised Learning114
Unsupervised Learning-Based Framework for Deepfake Video Detection114
Long Video Understanding With Learnable Retrieval in Video-Language Models113
Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames112
Transferable Backdoor Attack on Any CLIP Model With Any Target Class by Pre-Trained Hack Network111
Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude Similarity111
Asymptotics-Aware Multi-View Subspace Clustering110
Structured Graph Reasoning for Traffic Anomaly Detection108
Self-Guided Discriminative Locality Preserving Projections107
Vulnerability of Feature Extractors in 2D Image-Based 3D Object Retrieval105
Dynamic Mosaics: Saliency-Guided Adaptive Masking for Occluded Person Re-Identification104
MGKsite: Multi-Modal Knowledge-Driven Site Selection via Intra and Inter-Modal Graph Fusion104
Beyond Simple Extraction: Unleashing the Potential of Encoder Interaction in Few-Shot Segmentation103
Interpretable Graph Convolutional Network for Multi-View Semi-Supervised Learning102
GLCT: A Novel Global-Local Constraint for Unpaired Image-to-Image Translation100
TFBF: Temporal-Frequency Bidirectional Fusion for Action Quality Assessment100
Outliers Adaptation Exploration and Centroids Matching Label Refinement for Unsupervised Person Re-Identification100
Cross-modal Semantic Relevance is An Efficient Gatekeeper for Audio-Visual Video Parsing98
SkyML: A MLaaS Federation Design for Multicloud-Based Multimedia Analytics98
ICE: Interactive 3D Game Character Facial Editing via Dialogue97
Ensemble Prototype Networks for Unsupervised Cross-Modal Hashing With Cross-Task Consistency96
Siamese Alignment Network for Weakly Supervised Video Moment Retrieval96
MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition96
Anomaly-Led Prompting Learning Caption Generating Model and Benchmark96
Distilling Multi-View Diffusion Models Into 3D Generators94
Skeleton-Based Action Recognition With Select-Assemble-Normalize Graph Convolutional Networks94
Adversarial 3D-to-Real Watermarking: Revealing Invisible Messages Hidden in Complexly Distorted Surfaces93
BASNet: Boundary Assisted Network for Image Splicing Forgery Detection91
Scale Up Composed Image Retrieval Learning via Modification Text Generation91
Pixel Bleach Network for Detecting Face Forgery Under Compression90
Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image Segmentation90
3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature Enhancement88
$\rm {M}^{2}\rm {C-EvDet}$: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection88
Progressive Local Filter Pruning for Image Retrieval Acceleration87
XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework86
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation With Interpretability85
Semi-Supervised Domain Adaptation for Major Depressive Disorder Detection84
Feature First: Advancing Image-Text Retrieval Through Improved Visual Features83
Dynamic Contrastive Distillation for Image-Text Retrieval83
0.37651801109314