IEEE Transactions on Pattern Analysis and Machine Intelligence

Papers
(The TQCC of IEEE Transactions on Pattern Analysis and Machine Intelligence is 26. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Front Cover4106
Symbolic Visual Reinforcement Learning: A Scalable Framework With Object-Level Abstraction and Differentiable Expression Search2169
Editorial: Introduction to the Special Section on Best of CVPR'20221708
Implicit Annealing in Kernel Spaces: A Strongly Consistent Clustering Approach1120
Video Demoireing Using Focused-Defocused Dual-Camera System979
A Hybrid Stochastic-Deterministic Minibatch Proximal Gradient Method for Efficient Optimization and Generalization865
Towards Accurate and Compact Architectures via Neural Architecture Transformer855
Seeing Through Satellite Images at Street Views833
Next Bit Prediction: A Unified Lossless and Lossy Point Cloud Geometry Compression Framework828
Modeling Noisy Annotations for Point-Wise Supervision752
Test-Time Correction: An Online 3D Detection System via Visual Prompting707
Reliable and Compact Graph Fine-Tuning Via Graph Sparse Prompting667
Event-Based Photometric Bundle Adjustment654
Principal Uncertainty Quantification With Spatial Correlation for Image Restoration Problems649
Rethinking Link Prediction for Directed Graphs628
One-for-All: Towards Universal Domain Translation With a Single StyleGAN624
Rethinking Rotation-Invariant Recognition of Fine-Grained Shapes From the Perspective of Contour Points581
Revisiting Transformation Invariant Geometric Deep Learning: An Initial Representation Perspective578
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning577
Quadratic Matrix Factorization With Applications to Manifold Learning568
Weakly Supervised Semantic Segmentation via Box-Driven Masking and Filling Rate Shifting554
Vertical Layering of Quantized Neural Networks for Heterogeneous Inference541
Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing Interpretation539
A Personalized and Privacy-Preserving Federated Transformer Framework for Multilingual Sentiment Analysis521
MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation Segmentation506
Graph Convolutional Module for Temporal Action Localization in Videos490
Face Forgery Detection by 3D Decomposition and Composition Search488
Omni-Training: Bridging Pre-Training and Meta-Training for Few-Shot Learning486
DAQE: Enhancing the Quality of Compressed Images by Exploiting the Inherent Characteristic of Defocus476
Learning to Guide a Saturation-Based Theorem Prover475
Moment-Reenacting: Inverse Motion Degradation With Cross-Shutter Guidance454
SNI-SLAM++: Tightly-Coupled Semantic Neural Implicit SLAM454
Optimization-Based Post-Training Quantization With Bit-Split and Stitching453
Sparse-to-Dense Matching Network for Large-Scale LiDAR Point Cloud Registration452
Improving Viewpoint Robustness for Visual Recognition via Adversarial Training442
Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations438
Enhancing Representations Through Heterogeneous Self-Supervised Learning430
Adaptive Surface Normal Constraint for Geometric Estimation From Monocular Images429
A Clustering Validity Index With Multi-Granularity Fusion for Multiple Fuzzy Clustering Algorithms426
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting413
Prior Image Guided Snapshot Compressive Spectral Imaging408
Are Graph Convolutional Networks With Random Weights Feasible?393
DVIS++: Improved Decoupled Framework for Universal Video Segmentation383
Motion-Aware Dynamic Graph Neural Network for Video Compressive Sensing357
Probing Synergistic High-Order Interaction for Multi-Modal Image Fusion346
Label Hierarchy Transition: Delving Into Class Hierarchies to Enhance Deep Classifiers342
SCGT: Toward Scalable and Comprehensive Graph Transformer338
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition336
Towards Expressive Spectral-Temporal Graph Neural Networks for Time Series Forecasting329
Digging Into Uncertainty-Based Pseudo-Label for Robust Stereo Matching328
Asymmetric Convolution: An Efficient and Generalized Method to Fuse Feature Maps in Multiple Vision Tasks322
SymBOL: A General-Purpose Symbolic Learner for Scientific Discovery Using Bayesian Optimization-Enhanced Large Language Models317
Multi-Dataset, Multitask Learning of Egocentric Vision Tasks315
Learning Signed Hyper Surfaces for Oriented Point Cloud Normal Estimation315
Invariant Policy Learning: A Causal Perspective311
Simplicial Complex Neural Networks304
Locating and Counting Heads in Crowds With a Depth Prior304
OPAL: Occlusion Pattern Aware Loss for Unsupervised Light Field Disparity Estimation303
Affective Image Content Analysis: Two Decades Review and New Perspectives300
BiBBDM: Bidirectional Image Translation With Brownian Bridge Diffusion Models298
Interactive NeRF Geometry Editing With Shape Priors293
Guaranteed Tensor Recovery Fused Low-rankness and Smoothness290
Rein++: Efficient Generalization and Adaptation for Semantic Segmentation With Vision Foundation Models290
Metrics for Dataset Demographic Bias: A Case Study on Facial Expression Recognition289
VATr++: Choose Your Words Wisely for Handwritten Text Generation283
Learning Graph Convolutional Networks for Multi-Label Recognition and Applications283
Self-Supervised Skeleton Representation Learning Via Actionlet Contrast and Reconstruct279
Deep Long-Tailed Learning: A Survey277
Point-to-Pixel Prompting for Point Cloud Analysis With Pre-Trained Image Models273
ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting270
Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution262
Active Supervised Cross-Modal Retrieval257
Ensemble-Enhanced Semi-Supervised Learning With Optimized Graph Construction for High-Dimensional Data256
Learn to Predict Sets Using Feed-Forward Neural Networks254
Separable Spatial-Temporal Residual Graph for Cloth-Changing Group Re-Identification253
Task-Oriented Channel Attention for Fine-Grained Few-Shot Classification252
Face Generation and Editing With StyleGAN: A Survey251
Centerless Clustering248
Analysis of Video Quality Datasets via Design of Minimalistic Video Quality Models245
A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation244
On the Trade-Off Between Flatness and Optimization in Distributed Learning244
Transformer-Based Visual Segmentation: A Survey241
Physics-Informed Guided Disentanglement in Generative Networks240
Interaction-Based Inductive Bias in Graph Neural Networks: Enhancing Protein-Ligand Binding Affinity Predictions From 3D Structures239
Detection-Friendly Dehazing: Object Detection in Real-World Hazy Scenes239
Learning With Style: Continual Semantic Segmentation Across Tasks and Domains236
Structure-Preserving Image Super-Resolution235
AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic Domain229
S-NeRF++: Autonomous Driving Simulation via Neural Reconstruction and Generation225
Towards Unified Deep Image Deraining: A Survey and a New Benchmark224
Cover 2221
Rate-Distortion Theory in Coding for Machines and Its Applications216
Towards Pointsets Representation Learning via Self-Supervised Learning and Set Augmentation215
Human Interaction Understanding With Consistency-Aware Learning213
Learning Graph Attentions via Replicator Dynamics212
Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True and Class-Wise Distillation212
On Positive-Unlabeled Classification From Corrupted Data in GANs210
A Unified Experience Replay Framework for Spiking Deep Reinforcement Learning205
Boosting the Performance of Decentralized Federated Learning via Catalyst Acceleration200
Curriculum-Based Asymmetric Multi-Task Reinforcement Learning197
Influence Function Based Second-Order Channel Pruning: Evaluating True Loss Changes for Pruning is Possible Without Retraining197
Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning194
Reconstruction Guided Meta-Learning for Few Shot Open Set Recognition194
Compositional Scene Representation Learning via Reconstruction: A Survey190
Thermal3D-GS: Physics-Induced 3D Gaussians for Thermal Infrared Novel-View Synthesis With a Large-Scale Dataset189
Autonomous Causal Discovery: Evaluating LLMs' Priors and Constraint Strategies for Reliability188
MESA: Effective Matching Redundancy Reduction by Semantic Area Segmentation188
Scanpath Prediction in Panoramic Videos Via Expected Code Length Minimization187
Bridging Actions: Generate 3D Poses and Shapes In-Between Photos186
M$^{3}$3D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-Level Information Extraction185
Universal Image Segmentation With Efficiency179
SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face Intrinsics178
Ensemble Multi-Quantiles: Adaptively Flexible Distribution Prediction for Uncertainty Quantification178
Learning Efficient Meshflow and Optical Flow From Event Cameras178
LMP-GAN: Out-of-Distribution Detection for Non-Control Data Malware Attacks177
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis With Semantic Graph Prior177
Discriminant Feature Extraction by Generalized Difference Subspace177
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration176
Single-Stage Instance Segmentation Survey: A “1+2+3” Technical Framework175
Image-to-Image Translation With Disentangled Latent Vectors for Face Editing174
Supervised Small-Baseline and Large-Baseline Homography Learning With Diffusion-Based Data Generation173
Deciphering Object Concepts: Hierarchical Cross-Modal Relational Reasoning for Mining Object-Attribute-Affordance Associations172
Privacy Preserving Decentralized Learning With Positive-Incentive Noise172
MoIL: Momentum Imitation Learning for Efficient Vision-Language Adaptation171
GenPoly: Learning Generalized and Tessellated Shape Priors via 3D Polymorphic Evolving170
Physics-Informed Matrix Factorization Operator168
Graph-Oriented Instruction Tuning of Large Language Models for Generic Graph Mining168
GradMDM: Adversarial Attack on Dynamic Networks166
EVDI++: Event-based Video Deblurring and Interpolation via Self-Supervised Learning166
Adaptive Transfer Kernel Learning for Transfer Gaussian Process Regression164
A Unified Decision Rule for Generalized Out-of-Distribution Detection163
Reduced-Rank Tensor-on-Tensor Regression and Tensor-Variate Analysis of Variance163
SPARE: Symmetrized Point-to-Plane Distance for Robust Non-Rigid 3D Registration162
Human-Centric Transformer for Domain Adaptive Action Recognition155
Weakly Supervised Tracklet Association Learning With Video Labels for Person Re-Identification153
Image Lens Flare Removal Using Adversarial Curve Learning150
Learning to See Through With Events149
Knowledge-Based Embodied Question Answering148
Winsor-CAM: Human-Tunable Visual Explanations From Deep Networks via Layer-Wise Winsorization148
Learning From Partially Labeled Data for Multi-Organ and Tumor Segmentation147
GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector147
Matrix Completion via Non-Convex Relaxation and Adaptive Correlation Learning146
A New Brain Network Construction Paradigm for Brain Disorder via Diffusion-Based Graph Contrastive Learning146
Unbiased Scene Graph Generation via Two-Stage Causal Modeling145
To Fold or Not to Fold: Graph Regularized Tensor Train for Visual Data Completion145
SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation145
Self-Supervised Multimodal Learning: A Survey145
BNET: Batch Normalization With Enhanced Linear Transformation143
A Fully Automated Method for 3D Individual Tooth Identification and Segmentation in Dental CBCT141
Meta Invariance Defense Towards Generalizable Robustness to Unknown Adversarial Attacks141
Revisiting Transferable Adversarial Images: Systemization, Evaluation, and New Insights141
Evaluation for Weakly Supervised Object Localization: Protocol, Metrics, and Datasets140
Hypergraph-Based Multi-View Action Recognition Using Event Cameras140
Inter-Intra Hypergraph Computation for Survival Prediction on Whole Slide Images140
Deep Orientational Representation Learning for Ordinal Regression139
Understanding the Effects of Projectors in Knowledge Distillation139
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning139
Correcting Optical Aberration via Depth-Aware Point Spread Functions139
Mining Association Patterns From Neighborhood Insight139
QDTrack: Quasi-Dense Similarity Learning for Appearance-Only Multiple Object Tracking138
Differentially Private Graph Neural Networks for Whole-Graph Classification138
Asymmetric Loss Functions for Noise-Tolerant Learning: Theory and Applications138
Temporal Feature Matters: A Framework for Diffusion Model Quantization138
EAR-Net: Pursuing End-to-End Absolute Rotations from Multi-View Images138
Fear-Neuro-Inspired Reinforcement Learning for Safe Autonomous Driving136
HiGCIN: Hierarchical Graph-Based Cross Inference Network for Group Activity Recognition135
Deep Gait Recognition: A Survey135
Scalable Optimal Transport Methods in Machine Learning: A Contemporary Survey134
Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence Gap133
A Variational EM Acceleration for Efficient Clustering at Very Large Scales132
VNVC: A Versatile Neural Video Coding Framework for Efficient Human-Machine Vision132
PathNet: Path-Selective Point Cloud Denoising132
Variational Data-Free Knowledge Distillation for Continual Learning131
From Simple to Complex Scenes: Learning Robust Feature Representations for Accurate Human Parsing129
Controllable Generation With Text-to-Image Diffusion Models: A Survey129
Advances and Challenges in Meta-Learning: A Technical Review129
P2T: Pyramid Pooling Transformer for Scene Understanding128
Self-Scalable Tanh (Stan): Multi-Scale Solutions for Physics-Informed Neural Networks128
Deep Learning-Based Point Cloud Compression: An In-Depth Survey and Benchmark128
Accurate and Efficient Stereo Matching via Attention Concatenation Volume126
Revisiting Nonlocal Self-Similarity from Continuous Representation124
Random Permutation Set Reasoning123
MetaDrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement Learning122
Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses122
Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble122
AutoNovel: Automatically Discovering and Learning Novel Visual Categories122
Out-of-Domain Generalization From a Single Source: An Uncertainty Quantification Approach122
Robust Multimodal Learning With Missing Modalities via Parameter-Efficient Adaptation122
Flare7K++: Mixing Synthetic and Real Datasets for Nighttime Flare Removal and Beyond122
Enhancing Photorealism Enhancement121
Interpretable Optimization-Inspired Unfolding Network for Low-Light Image Enhancement120
Low-Shot Video Object Segmentation118
On the Robustness of Average Losses for Partial-Label Learning118
WildVideo: Benchmarking LMMs for Understanding Video-Language Interaction117
ComputingEdge ad117
Homeomorphism Prior for False Positive and Negative Problem in Medical Image Dense Contrastive Representation Learning116
Distributionally Location-Aware Transferable Adversarial Patches for Facial Images116
Cover 3116
luvHarris: A Practical Corner Detector for Event-Cameras116
RGB-T Tracking With Template-Bridged Search Interaction and Target-Preserved Template Updating114
Adversarially Robust Neural Architectures114
Modeling the Label Distributions for Weakly-Supervised Semantic Segmentation113
Compositional Physical Reasoning of Objects and Events From Videos112
Adaptive Perspective Distillation for Semantic Segmentation112
Cascaded Dynamic Memory Refinement and Semantic Alignment for Exo-to-Ego Cross-View Video Generation112
Semi-Supervised Learning for FGVC With Out-of-Category Data110
FreeFusion: Infrared and Visible Image Fusion via Cross Reconstruction Learning109
PMGT-VR: A Decentralized Proximal-Gradient Algorithmic Framework With Variance Reduction109
Unified Modality Separation: A Vision-Language Framework for Unsupervised Domain Adaptation109
Point Spatio-Temporal Transformer Networks for Point Cloud Video Modeling109
Pragmatic Communication in Multi-Agent Collaborative Perception107
Probabilistic Directed Distance Fields for Ray-Based Shape Representations107
DeepMesh: Differentiable Iso-Surface Extraction107
Human as Points: Explicit Point-Based 3D Human Reconstruction From Single-View RGB Images105
Any Fashion Attribute Editing: Dataset and Pretrained Models105
Hypergraph-Based High-Order Correlation Analysis for Large-Scale Long-Tailed Data Classification104
Deciphering the Feature Representation of Deep Neural Networks for High-Performance AI104
On the Universal Approximation Properties of Deep Neural Networks Using MAM Neurons103
YOTO++: Learning Long-Horizon Closed-Loop Bimanual Manipulation from One-Shot Human Video Demonstrations103
Supervision by Denoising102
Reframing Neural Networks: Deep Structure in Overcomplete Representations102
Pixel Distillation: Cost-Flexible Distillation Across Image Sizes and Heterogeneous Networks102
The Cluster Structure Function101
An Energy-Based Prior for Generative Saliency101
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding101
Reusable Architecture Growth for Continual Stereo Matching100
PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference100
Editorial: Special Section on Egocentric Perception100
Continual Unsupervised Generative Modeling100
SS-NeRF: Physically Based Sparse Spectral Rendering With Neural Radiance Field100
GLC++: Source-Free Universal Domain Adaptation Through Global-Local Clustering and Contrastive Affinity Learning99
Temporal Stereo Matching From Event Cameras via Joint Learning With Stereoscopic Flow99
Joint Framework for Single Image Reconstruction and Super-Resolution With an Event Camera99
Dynamic Differential Image Circle Diameter Measurement Precision Assessment: Application to Burning Droplets99
STAR-FC: Structure-Aware Face Clustering on Ultra-Large-Scale Graphs99
ModeRNN: Harnessing Spatiotemporal Mode Collapse in Unsupervised Predictive Learning98
Relationship Quantification of Image Degradations98
3D Visual Saliency: An Independent Perceptual Measure or a Derivative of 2D Image Saliency?98
Learn to Enhance Sparse Spike Streams98
Deep Learning on Object-Centric 3D Neural Fields98
S$^{2}$ 2O: Enhancing Adversarial Training With Second-Order Statistics of Weights97
MoBluRF: Motion Deblurring Neural Radiance Fields for Blurry Monocular Video97
Self-Guidance: Boosting Flow and Diffusion Generation on Their Own97
Compositional Generative Model of Unbounded 4D Cities97
An Algebraic Geometry Approach to Viewing Graph Solvability96
Stimulative Training++: Go Beyond the Performance Limits of Residual Networks96
AutoEval: Are Labels Always Necessary for Classifier Accuracy Evaluation?96
ONNXPruner: ONNX-Based General Model Pruning Adapter96
GhostingNet: A Novel Approach for Glass Surface Detection With Ghosting Cues95
CuDi: Curve Distillation for Efficient and Controllable Exposure Adjustment95
0.13801717758179