IEEE Transactions on Image Processing

Papers
(The TQCC of IEEE Transactions on Image Processing is 17. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Variational Structured Attention Networks for Deep Visual Representation Learning1091
TSFormer: Efficient Ultra-High-Definition Image Restoration via Trusted Min- p995
An Explanation Method Based on Interpretable Linear Model With Four Key Characteristics923
Density-Guided Incremental Dominant Instance Exploration for Two-View Geometric Model Fitting830
Color Spike Camera Reconstruction via Long Short-Term Temporal Aggregation of Spike Signals810
Equivariant Local Reference Frames With Optimization for Robust Non-Rigid Point Cloud Correspondence646
Star-Shaped Multi-Person Interaction Graph Model for Group Skeleton-Based Action Recognition382
Bi-Nuclear Tensor Schatten-p Norm Minimization for Multi-View Subspace Clustering324
AdaAugment: A Tuning-Free and Adaptive Approach to Enhance Data Augmentation321
Cross-Modality Pyramid Alignment for Visual Intention Understanding313
COME: A Collaborative Optimization Framework With Low-Rank MoE for Indoor 3D Object Detection295
High-Fidelity Seismic Super-Resolution Using Prior-Informed Deep Learning With 3D Awareness292
Zero-Pose-Prior NeRF: Recursive Radiance Field Reconstruction From Unposed and Unordered Images275
Advancing Pre-Trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection240
Consensus Sparsity: Multi-Context Sparse Image Representation via L -Induced Matrix Variate220
SemiRS-COC: Semi-Supervised Classification for Complex Remote Sensing Scenes With Cross-Object Consistency219
Leveraging Feature Alignment in Grassmannian Manifold for Multi-Output Regression Tasks212
One-Class Classification Using ℓp-Norm Multiple Kernel Fisher Null Approach211
Cross-Domain Few-Shot Medical Image Segmentation via Dynamic Semantic Matching205
Pro2Diff: Proposal Propagation for Multi-Object Tracking via the Diffusion Model201
Pose-Appearance Relational Modeling for Video Action Recognition197
Uncertainty-Guided Refinement for Fine-Grained Salient Object Detection182
Spatial Frequency Modulation Network for Efficient Image Dehazing180
Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning167
Toward Efficient Test Time Adaptation With Hierarchical Distribution Alignment164
MaCon: A Generic Self-Supervised Framework for Unsupervised Multimodal Change Detection155
FF-LPD: A Real-Time Frame-by-Frame License Plate Detector With Knowledge Distillation and Feature Propagation154
TTVFI: Learning Trajectory-Aware Transformer for Video Frame Interpolation153
Spectral State Fusion Tree Mamba for Hyperspectral Image Classification151
Language Supervised Multi-Camera Multi-Object Tracking148
Global Modeling Matters: A Fast, Lightweight, and Effective Baseline for Efficient Image Restoration144
An Adaptive Multi-Granularity Graph Representation of Image via Granular-ball Computing128
Toward Projected Clustering With Aggregated Mapping124
Cross-Modal Retrieval With Noisy Correspondence via Consistency Refining and Mining123
Focus on Finding Deepfakes: A Robust Proactive Detection Method Based on Orthogonal Moment Watermarking118
LearnMat: Semantic-Aware Self-Supervision Fine-Grained Visual Recognition117
Revisiting Fine-Grained Image Analysis by Semantic-Part Alignment114
H 3 Former: Hypergraph-Based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification113
Vision-Based UAV Self-Positioning in Low-Altitude Urban Environments112
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis From Language Models to Physics106
OccNeRF: Advancing 3D Occupancy Prediction in LiDAR-Free Environments105
STPNet: Scale-Aware Text Prompt Network for Medical Image Segmentation105
Graph Embedding Contrastive Multi-Modal Representation Learning for Clustering104
Automatic Quaternion-Domain Color Image Stitching103
HAda: Hyper-Adaptive Parameter-Efficient Learning for Multi-View ConvNets102
LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models101
Harnessing Multi-Modal Large Language Models for Measuring and Interpreting Color Differences98
Multi-Granularity Contrastive Cross-Modal Collaborative Generation for End-to-End Long-Term Video Question Answering98
Fine-Grained Recognition With Learnable Semantic Data Augmentation97
Attention-Guided Neural Networks for Full-Reference and No-Reference Audio-Visual Quality Assessment96
Fast 3D Room Layout Estimation Based on Compact High-Level Representation95
HAIMNet: A Hierarchical Adaptive Interaction Modulation Network for Low-Light Image Enhancement95
Generalization Beyond Feature Alignment: Concept Activation-Guided Contrastive Learning92
Perceptually Weighted Rate Distortion Optimization for Video-Based Point Cloud Compression91
Spatial-Temporal Scene Graph Generation for Open-Vocabulary Multiple Object Tracking89
Toward Generalizable Forgery Detection and Reasoning87
ScaleNet: Scaling up Pretrained Neural Networks With Incremental Parameters85
FD-SCU: Frequency Decomposition-Based Spectrum Collaborative Upsampling for Point Cloud Color Attribute85
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs83
Optimization-Inspired Learning With Architecture Augmentations and Control Mechanisms for Low-Level Vision83
MicroSDF: Microfacet-Driven Hybrid Neural SDFs for Mixed-Reflectance Surface Reconstruction82
SharpFormer: Learning Local Feature Preserving Global Representations for Image Deblurring82
MCIB: Multi-Modal Complementary Information Bottleneck for Hyperspectral and LiDAR Classification81
SRS: Siamese Reconstruction-Segmentation Network Based on Dynamic-Parameter Convolution81
Hyperspectral Meets Optical Flow: Spectral Flow Extraction for Hyperspectral Image Classification80
TSCCD: Temporal Self-Construction Cross-Domain Learning for Unsupervised Hyperspectral Change Detection80
Stacked Deconvolutional Network for Semantic Segmentation80
NeuralDiffuser: Neuroscience-Inspired Diffusion Guidance for fMRI Visual Reconstruction79
Weakly Supervised Semantic Segmentation via Alternate Self-Dual Teaching79
Rethinking Sampling Strategies for Unsupervised Person Re-Identification79
Point-Based Learnable Query Generator for Human–Object Interaction Detection78
Unsupervised Modality-Transferable Video Highlight Detection With Representation Activation Sequence Learning78
Toward Robust Alignment for Video Dehazing With Temporal Lookup Table78
Decoupled Cross-Modal Phrase-Attention Network for Image-Sentence Matching77
Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion Prediction76
SegHSI: Semantic Segmentation of Hyperspectral Images With Limited Labeled Pixels76
Boundary-Aware Prototype in Semi-Supervised Medical Image Segmentation76
Fuzzy Sparse Subspace Clustering for Infrared Image Segmentation75
Inverse Image Frequency for Long-Tailed Image Recognition74
Addressing Challenges of Incorporating Appearance Cues Into Heuristic Multi-Object Tracker via a Novel Feature Paradigm73
Interactive Face Video Coding: A Generative Compression Framework72
Cyclic Self-Training With Proposal Weight Modulation for Cross-Supervised Object Detection72
Toward Video Anomaly Retrieval From Video Anomaly Detection: New Benchmarks and Model72
Commonality Feature Representation Learning for Unsupervised Multimodal Change Detection72
ASDTracker: Adaptively Sparse Detection With Attention-Guided Refinement for Efficient Multi-Object Tracking71
Advances in Predictive RAHT for Geometric Point Cloud Compression71
Cross-Layer Contrastive Learning of Latent Semantics for Facial Expression Recognition71
Cross-Domain Diffusion With Progressive Alignment for Efficient Adaptive Retrieval71
Variational Bayes Image Restoration With Compressive Autoencoders71
Multi-Exposure Image Fusion via Deformable Self-Attention70
Transition Is a Process: Pair-to-Video Change Detection Networks for Very High Resolution Remote Sensing Images69
Motion and Appearance Decoupling Representation for Event Cameras69
FsaNet: Frequency Self-Attention for Semantic Segmentation68
Learning Dynamic Prompts for All-in-One Image Restoration68
Distractor-Aware Event-Based Tracking67
Precise Facial Landmark Detection by Reference Heatmap Transformer67
Non-Cascaded and Crosstalk-Free Multi-Image Encryption Based on Optical Scanning Holography Using 2D Orthogonal Compressive Sensing67
Soft Supervision-Guided Spatial–Temporal Refinement Network for Video-Based Visible-Infrared Person Re-Identification67
KSS-ICP: Point Cloud Registration Based on Kendall Shape Space67
RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation66
BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation66
NR-MVSNet: Learning Multi-View Stereo Based on Normal Consistency and Depth Refinement65
DCD-UIE: Decoupled Chromatic Diffusion Model for Underwater Image Enhancement65
Quality-Aware Spatio-Temporal Transformer Network for RGBT Tracking64
Inter-LPCM: Learning-Based Inter-Frame Predictive Coding for LiDAR Point Cloud Compression64
RoMo: Robust Unsupervised Multimodal Learning With Noisy Pseudo Labels64
Fine-Grained Spatio-Temporal Parsing Network for Action Quality Assessment64
HANeRV: Hierarchically Adaptive Neural Representation for Video Compression63
Semi-Negative Contrastive Subclass Discriminative Network for Compositional Zero-Shot Learning63
Energy-Based Domain Adaptation Without Intermediate Domain Dataset for Foggy Scene Segmentation63
Implicit-Explicit Integrated Representations for Multi-View Video Compression63
Partition Map Prediction for Fast Block Partitioning in VVC Intra-Frame Coding63
Bayesian Multifractal Image Segmentation62
Uncertainty Quantification for Semi-Supervised Object Detection in Remote Sensing Images62
UVaT: Uncertainty Incorporated View-Aware Transformer for Robust Multi-View Classification62
Counterfactual Risk Minimization for Out-of-Distribution Generalization61
Rich Action-Semantic Consistent Knowledge for Early Action Prediction61
HOPE: Enhanced Position Image Priors via High-Order Implicit Representations60
IMPRESS: Incomplete Human Motion Prediction via Motion Recovery and Structural-Semantic Fusion60
Learned Spherical Image Compression With Spherical Convolution-Self-Attention and Transformer Context Model60
LB-PTQ: Effective Low-Bit Post-Training Quantization for Vision Transformers60
A Discrete-Mapping-Based Cross-Component Prediction Paradigm for Screen Content Coding59
PolarPose: Single-Stage Multi-Person Pose Estimation in Polar Coordinates59
Generalizing to Out-of-Sample Degradations via Model Reprogramming59
Characteristic Mapping for Ellipse Detection Acceleration58
LNet: Lightweight Network for Driver Attention Estimation via Scene and Gaze Consistency58
Positional Encoding Image Prior57
Perception-Inspired Network for Stereo Image Quality Assessment57
Robust Ellipse Fitting Based on Maximum Correntropy Criterion With Variable Center56
Unsupervised Domain Adaptation in Biomedical Images Segmentation With Guided Diffusion Generative Prior56
Multispectral Snapshot Image Registration Using Learned Cross Spectral Disparity Estimation and a Deep Guided Occlusion Reconstruction Network56
Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization56
Causal Inference Hashing for Long-Tailed Image Retrieval55
Self-Anchored Progressive Framework With Noise Mitigation for Unsupervised Camouflaged Object Detection54
IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting54
Image Reconstruction for Accelerated MR Scan With Faster Fourier Convolutional Neural Networks54
Noise Prior Knowledge Informed Bayesian Inference Network for Hyperspectral Super-Resolution54
ReCoTR: Reducing Semantic Cognitive Shift via Dual-Consensus Token Compression for Remote Sensing Image-Text Retrieval54
RobustMat: Neural Diffusion for Street Landmark Patch Matching Under Challenging Environments53
Attack-Augmented Mixing-Contrastive Skeletal Representation Learning52
CGMNet: A Center-Pixel and Gated Mechanism-Based Attention Network for Hyperspectral Change Detection52
Shared Manifold Regularized Joint Feature Selection for Joint Classification and Regression in Alzheimer’s Disease Diagnosis51
Physics-Guided Cross-Modal Decoupling With Test-Time Adaptation for Hyperspectral Image Restoration51
Individual and Common Attack: Enhancing Transferability in VLP Models Through Modal Feature Exploitation51
Exploring the Potential of Pooling Techniques for Universal Image Restoration51
NesTD-Net: Deep NESTA-Inspired Unfolding Network With Dual-Path Deblocking Structure for Image Compressive Sensing51
Joint Denoising-Demosaicking Network for Long-Wave Infrared Division-of-Focal-Plane Polarization Images With Mixed Noise Level Estimation50
MaskFaceGAN: High-Resolution Face Editing With Masked GAN Latent Code Optimization50
Rethinking Object Saliency Ranking: A Novel Whole-Flow Processing Paradigm49
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation49
Toward Robust and Unconstrained Full Range of Rotation Head Pose Estimation48
A Real-Time Memory Updating Strategy for Unsupervised Person Re-Identification48
Mask-Guided Asymmetric Contrastive and Semantic Alignment for Unsupervised Person Re-Identification48
KeypointDiff: Keypoints-Guided Diffusion Model for Unpaired Object-Level SAR-to-Optical Aircraft Image Translation48
Reduced Biquaternion Dual-Branch Deraining U-Network via Multi-Attention Mechanism48
Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal48
Enhancing Few-Shot Out-of-Distribution Detection With Pre-Trained Model Features47
High-Quality and Diverse Few-Shot Image Generation via Masked Discrimination47
CAS-ViT: Convolutional Additive Self-Attention Vision Transformers for Efficient Mobile Applications47
Cross-Modal Causal Representation Learning for Radiology Report Generation47
Double Oracle Neural Architecture Search for Game Theoretic Deep Learning Models46
Neural Scene Designer: Self-Styled Semantic Image Manipulation46
Multi-Scale Fusion and Decomposition Network for Single Image Deraining46
Fast Learning Radiance Fields by Shooting Much Fewer Rays46
Reviewer Summary for Transactions on Image Processing46
Mutually Reinforcing Learning of Decoupled Degradation and Diffusion Enhancement for Unpaired Low-Light Image Lightening46
LPATR-Net: Learnable Piecewise Affine Transformation Regression Assisted Data-Driven Dehazing Framework46
UniEmoX: Cross-Modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception45
Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance Minimization45
MotionPrior: Exploring Efficient Learning of Motion Concepts for Few-Shot Video Generation45
SIR: Self-Supervised Image Rectification via Seeing the Same Scene From Multiple Different Lenses45
Unsupervised Domain Adaptive Object Detection via Semantic Consistency and Compactness Learning45
Single-Image Reflection Removal via Iterative Prompt Learning of Reflection Level44
Sensitivity Decouple Learning for Image Compression Artifacts Reduction44
U-N2C: A Dual Memory-Guided Disentanglement Framework for Unsupervised System Matrix Denoising in Magnetic Particle Imaging44
COMBINER: Composed Image Retrieval Guided by Attribute-Based Neighbor Relations44
Degraded Reference Image Quality Assessment43
PFONet: A Progressive Feedback Optimization Network for Lightweight Single Image Dehazing43
IEEE Transactions on Image Processing publication information43
SD-ReID: View-Aware Stable Diffusion for Aerial-Ground Person Re-Identification43
Image-Level Adaptive Adversarial Ranking for Person Re-Identification43
Two-Timer-KAN: Dual-Exclusive Fourier KANs With Gaussian Fusion for Few-Shot Multimodal Remote Sensing Imagery Classification43
Bidirectional Cross-Modal Collaborative Alignment via Semantic-Guided Visual Embeddings for Partially Relevant Video Retrieval43
FABNet: Frequency-Aware Binarized Network for Single Image Super-Resolution42
MA-ST3D: Motion Associated Self-Training for Unsupervised Domain Adaptation on 3D Object Detection42
One Step Diffusion-Based Super-Resolution With Time-Aware Distillation42
PointFormer: Keypoint-Guided Transformer for Simultaneous Nuclei Segmentation and Classification in Multi-Tissue Histology Images42
Interpretable Neural Networks for Video Separation: Deep Unfolding RPCA With Foreground Masking42
Long-Tailed and Inter-Class Homogeneity Matters in Multi-Class Weakly Supervised Tissue Segmentation of Histopathology Images42
Source-Guided Target Feature Reconstruction for Cross-Domain Classification and Detection42
Advancing Video Anomaly Detection: A Bi-Directional Hybrid Framework for Enhanced Single- and Multi-Task Approaches42
BP-NeRF: End-to-End Neural Radiance Fields for Sparse Images Without Camera Pose in Complex Scenes42
Perception-Guided Quality Metric of 3D Point Clouds Using Hybrid Strategy41
Enhancing Target Recognition Performance in SSVEP-Based Brain–Computer Interfaces via Deep Neural Networks With Pyramid Squeeze Attention41
DynSUP: Dynamic Gaussian Splatting From an Unposed Image Pair41
Dynamic Atomic Column Detection in Transmission Electron Microscopy Videos via Ridge Estimation41
Hierarchical Hashing Learning for Image Set Classification41
AirDC: Adaptive Iterative Depth Refinement Framework for Full-Range Metric Depth Completion41
Rotational Convolution: Rethinking Convolution for Downside Fisheye Images41
MBFQuant: A Multiplier-Bitwidth-Fixed, Mixed-Precision Quantization Method for Mobile CNN-Based Applications40
Learning Domain Invariant Representations for Generalizable Person Re-Identification40
Restoration of Images Taken Through a Dirty Window Using Optics-Guided Transformer40
Enhancing Text-Based Person Retrieval by Combining Fused Representation and Reciprocal Learning With Adaptive Loss Refinement40
Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense Prediction40
Hyperbolic Cycle Alignment for Infrared–Visible Image Fusion40
Bayesian Nonnegative Tensor Completion With Automatic Rank Determination40
Improving Unsupervised Ultrasonic Image Anomaly Detection via Frequency-Spatial Feature Filtering and Gaussian Mixture Modeling39
Zero-Shot Camouflaged Object Detection39
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks39
Dynamic Slimmable Denoising Network39
Cross-Modal Contrastive Learning Network for Few-Shot Action Recognition39
CKD: Contrastive Knowledge Distillation From a Sample-Wise Perspective39
DACESR: Degradation-Aware Conditional Embedding for Real-World Image Super-Resolution38
Deep Underwater Image Quality Assessment With Explicit Degradation Awareness Embedding38
Underwater Image Enhancement via Intelligent Optimized Multi-Exposure Image Fusion38
Learned Image Compression With Gaussian-Laplacian-Logistic Mixture Model and Concatenated Residual Modules38
Rethinking the Low-Light Video Enhancement: Benchmark Datasets and Methods38
Text-Driven Relation Manipulation of Diffusion Imagery38
BVSR-EvD: Blurry Video Space-Time Super-Resolution With Events via Diffusion Models38
TransDiff: Unsupervised Non-Line-of-Sight Imaging With Aperture-Limited Relay Surfaces38
Marine Saliency Segmenter: Object-Focused Conditional Diffusion With Region-Level Semantic Knowledge Distillation38
BPMTrack: Multi-Object Tracking With Detection Box Application Pattern Mining37
Probabilistic Embeddings With Evidence Learning and Refinement for Text–Video Retrieval37
MILES: Metric Learning with Expandable Subspace for Pre-Trained Model-Based Class-Incremental Learning37
PCE-GAN: A Generative Adversarial Network for Point Cloud Attribute Quality Enhancement Based on Optimal Transport37
Contrastive Conditional Latent Diffusion for Audio-Visual Segmentation37
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation37
Multi-Label Auroral Image Classification Based on CNN and Transformer37
Learning Transferable Conceptual Prototypes for Interpretable Unsupervised Domain Adaptation37
TTST: A Top-k Token Selective Transformer for Remote Sensing Image Super-Resolution37
Hyperspectral Image Classification via Cascaded Spatial Cross-Attention Network37
SDSFusion: A Semantic-Aware Infrared and Visible Image Fusion Network for Degraded Scenes37
Magi-Net: Meta Negative Network for Early Activity Prediction36
Learning Structure Aware Deep Spectral Embedding36
Compression-Oriented Video Super-Resolution36
DVMark: A Deep Multiscale Framework for Video Watermarking36
Padé Neurons for Efficient Neural Models36
U-Shape Transformer for Underwater Image Enhancement36
Degradation-Adaptive Denoising: Aligning Diffusion Models With Physics of Video Snapshot Compressive Imaging35
Diverse Target and Contribution Scheduling for Domain Generalization35
Hyperpixels: Flexible 4D Over-Segmentation for Dense and Sparse Light Fields35
C-NeRF: Representing Scene Changes as Directional Consistency Difference-Based NeRF35
Learn From Examples: In-Context Learning for Camouflaged Object Detection35
DisAVR: Disentangled Adaptive Visual Reasoning Network for Diagram Question Answering35
Toward Transparent Deep Image Aesthetics Assessment With Tag-Based Content Descriptors35
Decoupling Discriminative Attributes for Few-Shot Fine-Grained Recognition35
Rethinking Generalized Zero-Shot Learning: A Synthesized Per-Instance Attribute Perspective35
Broadcast-Gated Attention With Identity Adaptive Integration for Efficient Image Super-Resolution35
Versatile Denoising-Based Approximate Message Passing for Compressive Sensing34
SUIT: Spatial-Spectral Union-Intersection Interaction Network for Hyperspectral Object Tracking34
0.59186911582947