ACM Transactions on Multimedia Computing Communications and Applicatio

Papers
(The TQCC of ACM Transactions on Multimedia Computing Communications and Applicatio is 9. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Hypercube Pooling for Visual Semantic Embedding554
Self-Adaptive Representation Learning Model for Multi-Modal Sentiment and Sarcasm Joint Analysis175
Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image173
ForgeFinder: Perceptive Multimodal Deepfake Detection via Multi-grained Forgery Localization127
Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action Localization113
Fine-Grained Text-to-Video Temporal Grounding from Coarse Boundary91
Unsupervised Discovery and Manipulation of Continuous Disentangled Factors of Variation87
Semi-supervised Learning for Mars Imagery Classification and Segmentation86
AED-PADA: Improving Generalizability of Adversarial Example Detection via Principal Adversarial Domain Adaptation86
Enhanced Video Super-Resolution Network towards Compressed Data77
Psychology-Guided Environment Aware Network for Discovering Social Interaction Groups from Videos70
Towards Intelligent Attack Detection Using DNA Computing70
Establishing Trust and Security in Decentralized Metaverse: A Web 3.0 Approach69
CVLP-NaVD: Contrastive Visual-language Pre-training Models for Non-annotated Visual Description64
Upsampling Algorithm for V-PCC-Coded 3D Point Clouds61
A Siamese Inverted Residuals Network Image Steganalysis Scheme based on Deep Learning59
QuickCSGModeling: Quick CSG Operations Based on Fusing Signed Distance Fields for VR Modeling59
Tensorial Evolutionary Optimization for Natural Image Matting57
Exploring Talking Head Models with Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation56
JDAN: Joint Detection and Association Network for Real-Time Online Multi-Object Tracking55
High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream Transformers55
Image Cropping with Content and Composition Attribute-aware Global Relation Reasoning52
Reconstruction-Free Image Compression for Machine Vision via Knowledge Transfer52
Attentional Composition Networks for Long-Tailed Human Action Recognition51
A Comprehensive Survey on Methods for Image Integrity51
Backdoor Two-Stream Video Models on Federated Learning49
LAM-MD: A Large Action Model Integrated Multi-Agent Reinforcement Learning Framework for NOMA Networks48
Towards Generalizable Deepfake Detection by Primary Region Regularization48
Disentangled Multimodal Tuning and Interaction for Human Perception Understanding48
Quantum Fourier Convolutional Network47
BiC-Net: Learning Efficient Spatio-temporal Relation for Text-Video Retrieval47
Infrared and Visible Image Fusion via Text-Prior Guided Frequency-Domain Decomposition47
SEADUNet: A Multilingual Ancient Document Image Binarization using EMCAM Attention Mechanism and SCP47
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation44
Joint Mixing Data Augmentation for Skeleton-Based Action Recognition44
Seeing in the Dark with Ambient Guidance44
A novel image hashing with visual and structural feature maps for content authentication43
HTTP Adaptive Streaming: A Review on Current Advances and Future Challenges43
Point Cloud Quality Assessment: Dataset Construction and Learning-based No-reference Metric42
New Metrics and Dataset for Biological Development Video Generation41
Expanding-Window Zigzag Decodable Fountain Codes for Scalable Multimedia Transmission40
Visual-linguistic-stylistic Triple Reward for Cross-lingual Image Captioning40
CLOUD-CODEC : A New Way of Storing Traffic Camera Footage at Scale39
Exploiting Instance-level Relationships in Weakly Supervised Text-to-Video Retrieval37
A Self-Defense Copyright Protection Scheme for NFT Image Art Based on Information Embedding37
Domain-Aware Semantic Alignment Hashing for Large-Scale Zero-Shot Image Retrieval37
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval36
DALD-PCAC: Density-Adaptive Learning Descriptor for Point Cloud Lossless Attribute Compression36
Boosting Transferability of Adversarial Examples with Spatio-Temporal Context36
Light Field Reconstruction Using Multi-orientation Epipolar Plane Images34
Universal Relocalizer for Weakly Supervised Referring Expression Grounding34
ViCoFace: Learning Disentangled Latent Motion Representations for Visual-Consistent Face Reenactment33
VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural Heritage33
HCMS: Hierarchical and Conditional Modality Selection for Efficient Video Recognition32
Robust Video Stabilization based on Motion Decomposition32
A Multi-Task Adversarial Attack against Face Authentication32
(Compress and Restore) N : A Robust Defense Against Adversarial Attacks on Image Classification31
Learned Image Compression with Frequency Feature Interaction and Non-local Cross-similarity Prior31
Decoupling Deep Learning for Enhanced Image Recognition Interpretability30
EiMOL: A Secure Medical Image Encryption Algorithm based on Optimization and the Lorenz System30
Immersive Multimedia Service Caching in Edge Cloud with Renewable Energy30
GANonymization: A GAN-Based Face Anonymization Framework for Preserving Emotional Expressions29
Detection of Moving Object Using Superpixel Fusion Network29
ER-Depth: Enhancing the Robustness of Self-Supervised Monocular Depth Estimation in Challenging Scenes29
Using Four Hypothesis Probability Estimators for CABAC in Versatile Video Coding29
SNIPPET: A Framework for Subjective Evaluation of Visual Explanations Applied to DeepFake Detection29
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding28
Boundary Attention-Guided Sparse Feature Learning for Underwater Object Tracking in Edge Computing28
Multi-spectral Class Center Network for Face Manipulation Localization27
Image Defogging Based on Regional Gradient Constrained Prior27
The Price of Unlearning: Identifying Unlearning Risk in Edge Computing26
A Quality of Experience and Visual Attention Evaluation for 360° Videos with Non-spatial and Spatial Audio26
GMS-3DQA: Projection-Based Grid Mini-patch Sampling for 3D Model Quality Assessment26
Source Information-Assisted UV-Space Transformation Network for Person Image Generation25
Reversible Data Hiding in Shared JPEG Images25
Human Selective Matting25
Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model25
Learning Speaker-Invariant Visual Features for Lipreading25
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation25
MDRA: A Motion-guided Dual-stream Recurrent Attention Framework for Dynamic Hand Gesture Recognition25
DTSD: A Dual Teacher–Student-Based Discrimination Model for Anomaly Detection25
Multi-Grained Point Cloud Geometry Compression via Dual-Model Prediction with Extended Octree25
Gleaning Wisdom from the Past: Towards Label Incremental Learning for Online Hashing with a Plug-and-Play Framework25
TEVL: Trilinear Encoder for Video-language Representation Learning24
One-Bit Supervision for Image Classification: Problem, Solution, and Beyond24
Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical Transformers24
Principal Component Approximation Network for Image Compression24
Deep Chroma Compression of Tone-Mapped Images24
SkiTrack: An Aerial Skiing Benchmark for Human-Centric Object Tracking24
Enhancing Embedding Diversity and Robustness for Image-Text Retrieval in Remote Sensing24
Zero-shot Scene Graph Generation via Triplet Calibration and Reduction23
Cyclic Self-attention for Point Cloud Recognition23
ATMNet: Adaptive Texture Migration Network for Guided Depth Super-Resolution23
Melody Generation from Lyrics with Local Interpretability23
Adversarial Sample Synthesis for Visual Question Answering22
Temporal and Semantic Correlation Network for Weakly-Supervised Temporal Action Localization22
DATRA-MIV: Decoder-Adaptive Tiling and Rate Allocation for MPEG Immersive Video22
Boosting Targeted Adversarial Transferability with Feature Contrastive Optimization22
Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-Identification22
An Efficient and Accurate GPU-based Deep Learning Model for Multimedia Recommendation22
QoE Evaluation for VR with Vibrotactile Feedback Based on Inter-user Brain Spatial Information22
Cross-modal Semantically Augmented Network for Image-text Matching22
Toward Egocentric Compositional Action Anticipation with Adaptive Semantic Debiasing22
Domain Adaptation for Cross-View Localization via Multi-Teacher Knowledge Distillation21
Dual Alignment-Enhanced Fashion Vision-Language Pre-Training21
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-Modal Sarcasm Detection21
THMM-CLIP: Task-Guided Hierarchical Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning21
Temporal Dynamic Concept Modeling Network for Explainable Video Event Recognition21
Hyperbolic Active Learning for Label-Efficient Action Segmentation21
Gloss-driven Conditional Diffusion Models for Sign Language Production21
Spotting the Fakes: A Deep Dive into GAN-Generated Face Detection21
QG-STR: Training-Time Optimized Question-Guided Scene Text Recognition via Visual Question Answering21
Learning Domain Invariant Features for Unsupervised Indoor Depth Estimation Adaptation20
DISA: Disentangled Dual-Branch Framework for Affordance-Aware Human Insertion20
Visual Security Index Combining CNN and Filter for Perceptually Encrypted Light Field Images20
LayoutEnc: Leveraging Enhanced Layout Representations for Transformer-based Complex Scene Synthesis20
Counterfactual Scenario-relevant Knowledge-enriched Multi-modal Emotion Reasoning20
Diversity-Representativeness Replay and Knowledge Alignment for Lifelong Vehicle Re-identification20
Text-Guided Synthesis of Masked Face Images19
ReFID: Reciprocal Frequency-aware Generalizable Person Re-identification via Decomposition and Filtering19
PTHUMAN3D: 3D Gaussian Human Avatar Modeling with the Poincaré Ball and the Triplane Representation19
Mutually-Guided Hierarchical Multi-Modal Feature Learning for Referring Image Segmentation19
Multiply Complementary Priors for Image Compressive Sensing Reconstruction in Impulsive Noise19
Structure-aware Video Style Transfer with Map Art18
CVAF: A CLIP-Based View-Consistent Alignment Framework for Aerial-Ground Person Re-Identification18
Multi-Task–Driven Adapter-Based Foundation Model for Locomotion Prediction in Virtual Reality18
StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition18
A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach18
SeGDP: Source-free Cross-domain Few-shot Learning via Semantic Guided Diversity Prompting18
Multigranularity Feature Aggregation and Cross-level Boundary Modeling for Temporal Action Detection18
Generative Image Steganography Based on Guidance Feature Distribution18
Cross-Modality Relation and Uncertainty Exploration for Text-Based Person Search18
Attack-Defending Contrastive Learning for Volumetric Medical Image Zero-Watermarking18
User-Generated Content and Editors in Games: A Comprehensive Survey18
Shot Boundary Detection Using Color Clustering and Attention Mechanism18
Triplet Contrastive Representation Learning for Unsupervised Vehicle Re-Identification17
Maximizing Long-Term Task Completion Ratio of UAV-Enabled Wirelessly Powered MEC Systems17
3D Facial Shape Similarity with Deep Perceptual Representations17
Robust RGB-T Tracking via Adaptive Modality Weight Correlation Filters and Cross-modality Learning17
Hyperbolic-Based Cross-Modal Semantic Remodeling Network for Zero-Shot Sketch-Based Image Retrieval17
CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding17
Query-Guided Prototype Learning with Decoder Alignment and Dynamic Fusion in Few-Shot Segmentation17
Sentiment-Oriented Transformer-Based Variational Autoencoder Network for Live Video Commenting17
Deep Modular Co-Attention Shifting Network for Multimodal Sentiment Analysis17
PADVG: A Simple Baseline of Active Protection for Audio-Driven Video Generation17
Potential Features Fusion Network for Multimodal Fake News Detection17
Multi-view Shape Generation for a 3D Human-like Body17
NSDIE: Noise Suppressing Dark Image Enhancement Using Multiscale Retinex and Low-Rank Minimization16
A Normalized Slicing-assigned Virtualization Method for 6G-based Wireless Communication Systems16
Arbitrary Virtual Try-on Network: Characteristics Preservation and Tradeoff between Body and Clothing16
Temporal Scene Montage for Self-Supervised Video Scene Boundary Detection16
Generating Robust Adversarial Examples against Online Social Networks (OSNs)16
Robust and Secure Hashing Towards Pirated Neural Network Model Detection16
Progressive Transformer Machine for Natural Character Reenactment16
Dual Dynamic Threshold Adjustment Strategy16
Quality Enhancement of Compressed 360-Degree Videos Using Viewport-based Deep Neural Networks16
Multi-Modal Driven Pose-Controllable Talking Head Generation16
Privacy-preserving Multi-source Cross-domain Recommendation Based on Knowledge Graph16
ProposalVLAD with Proposal-Intra Exploring for Temporal Action Proposal Generation16
PrivaMod: Uncertainty-Aware Multimedia Fusion with Privacy Guarantees for NFT Visual and Transaction Analysis16
Self-supervised Multi-view Learning via Auto-encoding 3D Transformations16
Domain-invariant and Patch-discriminative Feature Learning for General Deepfake Detection16
Cascaded Adaptive Graph Representation Learning for Image Copy-Move Forgery Detection16
Toward High-quality Face-Mask Occluded Restoration16
Generation and Editing of Mandrill Faces: Application to Sex Editing and Assessment16
3DMambaComplete: Structured State Space Model for High-Efficiency Point Cloud Completion16
Unsupervised Domain Adaptation by Causal Learning for Biometric Signal-based HCI16
GAN-Assisted Road Segmentation from Satellite Imagery15
Quality Assessment in the Era of Large Models: A Survey15
Learning the User’s Deeper Preferences for Multi-modal Recommendation Systems15
Semantics and Non-fungible Tokens for Copyright Management on the Metaverse and Beyond15
Skeleton-Aware Graph-Based Adversarial Networks for Human Pose Estimation from Sparse IMUs15
Learning Nighttime Semantic Segmentation the Hard Way15
Deep Relational Knowledge Distillation Hashing via Relaxed Masking Triplet Optimization for Large-scale Image Retrieval15
Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications15
Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper15
Offloading-based Power-Efficient Mobile VTuber Live Streaming15
Semantic Completion and Filtration for Image–Text Retrieval15
PingTactics: A Multimodal Dataset for Table Tennis Action Recognition and Tactical Analysis15
Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video Captioning15
SSAT: Active Authorization Control and User’s Fingerprint Tracking Framework for DNN IP Protection15
Autoregressive GAN for Semantic Unconditional Head Motion Generation14
Transformer-Based Visual Grounding with Cross-Modality Interaction14
Joint Structure-Texture Scan-Order for Point Cloud Attribute Compression Using Affine Transformation14
Portrait Video Compression with Semantic-guided Animation Model and Background Incremental Coding14
Dual Scene Graph Convolutional Network for Motivation Prediction14
Low-Latency Multimedia Delivery via Collaborative Cloud–Edge Caching in Edge Computing Networks14
Trans-Convo-Former Net for Hierarchical Prediction of Household Images14
Content-Aware Selective Encryption for H.265/HEVC Using Deep Hashing Network and Steganography14
Robust Image Hashing via CP Decomposition and DCT for Copy Detection14
FAST: Flexibly Controllable Arbitrary Style Transfer via Latent Diffusion Models14
A Collaborative Hierarchical Aggregation Network for Weakly Supervised Temporal Action Localization14
EIN: Exposure-Induced Network for Single-Image HDR Reconstruction14
Robust Long-Term Tracking via Localizing Occluders14
Self-supervised Calorie-aware Heterogeneous Graph Networks for Food Recommendation14
VRVul-Discovery: BiLSTM-based Vulnerability Discovery for Virtual Reality Devices in Metaverse14
Multiscale Feature Importance-Based Bit Allocation for End-to-End Feature Coding for Machines14
Full-body Human Motion Reconstruction with Sparse Joint Tracking Using Flexible Sensors14
Geometry-Insensitive RPN Prototypes for Domain Adaptive 3D Object Detection13
Generalizing Vision-Language Models to Novel Domains: A Comprehensive Survey13
Noise-Resistance Learning via Multi-Granularity Consistency for Unsupervised Domain Adaptive Person Re-Identification13
Language-guided Visual Tracking: Comprehensive and Effective Multimodal Information Fusion13
Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization13
Balanced and Accurate Pseudo-Labels for Semi-Supervised Image Classification13
Efficient Privacy-Preserving Video Analytics via Share Transforming in Distributed Clouds13
GJFusion: A Channel-Level Correlation Construction Method for Multimodal Physiological Signal Fusion13
Review and Analysis of RGBT Single Object Tracking Methods: A Fusion Perspective13
EVASR: Edge-Based Salience-Aware Super-Resolution for Enhanced Video Quality and Power Efficiency13
Language-guided Bias Generation Contrastive Strategy for Visual Question Answering13
Enhancing Pose-Guided Human Image Generation with Comprehensive and Adjustable 3D Control13
Generating and Evaluating Data of Daily Activities with an Autonomous Agent in a Virtual Smart Home13
InteractNet: Social Interaction Recognition for Semantic-rich Videos13
A Real-Time Medical Image Encryption Algorithm Leveraging a Novel Hypersensitive Chaotic Map13
Dynamic Weighted Gradient Reversal Network for Visible-infrared Person Re-identification13
Language-guided Residual Graph Attention Network and Data Augmentation for Visual Grounding13
A Simple Switchable Framework for Open-Vocabulary Video Instance Segmentation13
Boosting Few-shot Object Detection with Discriminative Representation and Class Margin13
Variational Autoencoder with CCA for Audio–Visual Cross-modal Retrieval12
iDAM: Iteratively Trained Deep In-loop Filter with Adaptive Model Selection12
CAQoE: A Novel No-Reference Context-aware Speech Quality Prediction Metric12
CALICE: Continuous Bitrate Control with Adapted LIC Model12
LFIZW-GRHFMR: Robust Zero-Watermarking with GRHFMR for Light Field Image12
Beyond the Parts: Learning Coarse-to-Fine Adaptive Alignment Representation for Person Search12
Complementary Feature Pyramid Network for Object Detection12
AMC: Adaptive Multi-expert Collaborative Network for Text-guided Image Retrieval12
A Review of Player Engagement Estimation in Video Games: Challenges and Opportunities12
Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation Learning12
Compressed Point Cloud Quality Index by Combining Global Appearance and Local Details12
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection12
T2C: Text-guided 4D Cloth Generation12
A Multimodal Hierarchical Attentional Ordering Network12
DPDFormer: A Coarse-to-Fine Model for Monocular Depth Estimation12
Instance-level Adversarial Source-free Domain Adaptive Person Re-identification12
How to Understand Named Entities: Using Commonsense for News Captioning12
Random Dense Knowledge Distillation for Continual Learning12
FishFormer: Annulus Slicing-based Transformer for Fisheye Rectification12
Dual-Modality-Shared Learning and Label Refinement for Unsupervised Visible-Infrared Person ReID12
Deep Differential Lifelong Cross-modal Hashing for Stream Medical Data Retrieval12
Unsupervised Adversarial Example Detection of Vision Transformers for Trustworthy Edge Computing12
LVT: A Learned Video Transcoding Framework12
ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic Segmentation12
Learning Semantic Representation on Visual Attribute Graph for Person Re-identification and Beyond12
A Hierarchically Discriminative Loss with Group Regularization for Fine-Grained Image Classification12
Context-Based Novel Histogram Bin Stretching Algorithm for Automatic Contrast Enhancement12
Boolean-based Two-in-One Secret Image Sharing by Adaptive Pixel Grouping12
Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing Modalities12
Learning to Discern Fine-Grained Cues across Domains: Generalizing ReID via Multi-Level Feature Propagation12
Adaptive Co-Operative Prompting and Uncertainty-Aware Implicit Knowledge Enhancement for Cross-Modal Retrieval11
Multi-Anchor Offset Representation Based Coarse-to-Fine Diffusion Model for Human Pose Estimation11
Style-FG: A Style-based Framework for Film Grain Analysis and Synthesis11
0.54181599617004