International Journal of Computer Vision

Papers
(The median citation count of International Journal of Computer Vision is 4. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Exploring the Semi-Supervised Video Object Segmentation Problem from a Cyclic Perspective980
Guest Editorial: Special Issue on Open-World Visual Recognition468
Bootstrapping Vision-Language Models for Frequency-Centric Self-Supervised Remote Physiological Measurement444
Common Pole–Polar Properties of Central Catadioptric Sphere and Line Images Used for Camera Calibration397
GenKL: An Iterative Framework for Resolving Label Ambiguity and Label Non-conformity in Web Images Via a New Generalized KL Divergence361
Guest Editorial: Special Issue on Large-Scale Generative Models for Content Creation and Manipulation239
MoDA: Modeling Deformable 3D Objects from Casual Videos214
Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting213
Robust Averaging using Adaptive Annealing202
Exocentric-to-Egocentric Adaptation for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs200
AutoIT: Automated Image Tagging with Random Perturbation196
Correction: Multi-source-free Domain Adaptive Object Detection195
Invert Your Prompt: Editing-Aware Diffusion Inversion176
Image-based Morphological Characterization of Filamentous Biological Structures with Non-constant Curvature Shape Feature172
Learning Extensible Series-Parallel Lookup Tables for Efficient Image Super-Resolution169
Large-Scale Pre-Trained Models Empowering Phrase Generalization in Temporal Sentence Localization169
View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements162
Instance-dependent Label Distribution Estimation for Learning with Label Noise147
SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels135
Correction: Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization134
A Minimal Solution for Image-Based Sphere Estimation132
PanAf20K: A Large Video Dataset for Wild Ape Detection and Behaviour Recognition131
Learning Text-to-Video Retrieval from Image Captioning130
Learning with Enriched Inductive Biases for Vision-Language Models129
Learning Accurate Performance Predictors for Ultrafast Automated Model Compression129
Conditional Temporal Variational AutoEncoder for Action Video Prediction127
From Open Set to Closed Set: Supervised Spatial Divide-and-Conquer for Object Counting125
RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion123
BioDrone: A Bionic Drone-Based Single Object Tracking Benchmark for Robust Vision114
EAN: Event Adaptive Network for Enhanced Action Recognition111
UniAttack: Unified Physical-Digital Face Attack Detection103
Weakly Supervised Salient Object Detection with Text Supervision99
Image Synthesis Under Limited Data: A Survey and Taxonomy97
Learning Discriminative Features for Visual Tracking via Scenario Decoupling94
MPANet: Motion Pattern Aggregation Network for Gait Recognition93
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks87
Delving Deeper into Anti-Aliasing in ConvNets87
FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention86
OpenMonkeyChallenge: Dataset and Benchmark Challenges for Pose Estimation of Non-human Primates85
Are Vision Transformers Robust to Spurious Correlations?80
Guest Editorial: Special Issue on the British Machine Vision Conference 202279
NAFT and SynthStab: A RAFT-Based Network and a Synthetic Dataset for Digital Video Stabilization78
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild78
Relating View Directions of Complementary-View Mobile Cameras via the Human Shadow77
In the Eye of Transformer: Global–Local Correlation for Egocentric Gaze Estimation and Beyond76
Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions76
Learning Latent Part-Whole Hierarchies for Point Clouds73
Bi-calibration Networks for Weakly-Supervised Video Representation Learning71
Correction: Consistent Prompt Tuning for Generalized Category Discovery71
Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization68
Learning Accurate Low-bit Quantization towards Efficient Computational Imaging68
Correction: SOTVerse: A User-Defined Task Space of Single Object Tracking67
CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database66
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation64
Symmetria: A Synthetic Dataset for Learning in Point Clouds61
Dynamic Knowledge Transfer for Mitigating Spurious Correlations in Deep Learning60
Skeleton Ground Truth Extraction: Methodology, Annotation Tool and Benchmarks58
UMSCS: A Novel Unpaired Multimodal Image Segmentation Method Via Cross-Modality Generative and Semi-supervised Learning58
Sample-efficient Audio-Visual Learning of Scene Acoustics57
Diagram Perception Networks for Textbook Question Answering via Joint Optimization57
AI killed the Video Star. Audio-Driven Diffusion Model for Expressive Talking Head Generation57
Weakly Supervised Training of Universal Visual Concepts for Multi-domain Semantic Segmentation56
Guest Editorial: Special Issue on the Promises and Dangers of Large Vision Models56
Project to Adapt: Domain Adaptation for Depth Completion from Noisy and Sparse Sensor Data55
Semantic-Based Implicit Feature Transform for Few-Shot Classification54
Deep Learning-Based Object Pose Estimation: A Comprehensive Survey54
Lightweight and Progressively-Scalable Networks for Semantic Segmentation53
UniCanvas: Affordance-Aware Unified Real Image Editing via Customized Text-to-Image Generation53
Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking50
A Motion-Based Compression and Tracking System for Video Camera Trap-Based Insect Behaviour Studies50
VideoQA in the Era of LLMs: An Empirical Study50
SeaFormer++: Squeeze-Enhanced Axial Transformer for Mobile Visual Recognition49
Exploiting Inter-Sample Affinity for Knowability-Aware Universal Domain Adaptation49
Sfnet: Faster and Accurate Semantic Segmentation Via Semantic Flow48
Free-view Face Relighting Using a Hybrid Parametric Neural Model on a SMALL-OLAT Dataset48
Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking47
Noise-Resistant Multimodal Transformer for Emotion Recognition47
Contrastive Mean Teacher for Robust Low-Light Image Enhancement46
BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher’s Bias45
UIL-AQA: Uncertainty-Aware Clip-Level Interpretable Action Quality Assessment44
Feature Hallucination for Self-supervised Action Recognition44
Cascaded Iterative Transformer for Jointly Predicting Facial Landmark, Occlusion Probability and Head Pose44
ICEv2: Interpretability, Comprehensiveness, and Explainability in Vision Transformer44
A Realism Metric for Generated LiDAR Point Clouds43
SRConvNet: A Transformer-Style ConvNet for Lightweight Image Super-Resolution43
Generative Adversarial Network Applications in Industry 4.0: A Review42
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts42
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates41
Basis Restricted Elastic Shape Analysis on the Space of Unregistered Surfaces41
Robust Partial-to-Partial Point Cloud Registration with Overlapping Mask Learning41
IEBins: Iterative Elastic Bins for Monocular Depth Estimation and Completion41
Improving Domain Adaptation Through Class Aware Frequency Transformation41
Paragraph-to-Image Generation with Information-Enriched Diffusion Model41
Understanding Synonymous Referring Expressions via Contrastive Features41
A Nonlinear, Regularized, and Data-independent Modulation for Continuously Interactive Image Processing Network41
A Generalized Contour Vibration Model for Building Extraction40
Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation40
Skeletonizing Caenorhabditis elegans Based on U-Net Architectures Trained with a Multi-worm Low-Resolution Synthetic Dataset40
Image Matting and 3D Reconstruction in One Loop40
Focal Modulation for Image Restoration39
Beyond Learned Metadata-Based Raw Image Reconstruction39
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era38
Watching Swarm Dynamics from Above: A Framework for Advanced Object Tracking in Drone Videos38
Modeling Scattering Effect for Under-Display Camera Image Restoration38
Relaxed Knowledge Distillation38
Cyclic Refiner: Object-Aware Temporal Representation Learning for Multi-view 3D Detection and Tracking38
From Forest to Zoo: Great Ape Behavior Recognition with ChimpBehave38
A CNN Based Approach for the Point-Light Photometric Stereo Problem37
Advances in 3D Neural Stylization: A Survey37
EfficientDeRain+: Learning Uncertainty-Aware Filtering via RainMix Augmentation for High-Efficiency Deraining36
Towards Fine-Grained Optimal 3D Face Dense Registration: An Iterative Dividing and Diffusing Method36
Control Color: Multimodal Diffusion-Based Interactive Image Colorization36
Correction: BaboonLand Dataset: Tracking Primates in the Wild and Automating Behaviour Recognition from Drone Videos34
Text2Scenes: Language-Guided Synthesis of Complex Indoor Scenes34
Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-identification34
Globally Correlation-Aware Hard Negative Generation34
Bigger Isn’t Always Better: Towards a General Prior for Medical Image Reconstruction34
WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation34
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models34
MedSegFM: A Generative Perspective for Lesion Segmentation via Flow Matching34
Guest Editorial: Special Issue on Computer Vision from 2D to 3D33
Robust Unpaired Image Dehazing via Density and Depth Decomposition33
I2DFormer+: Learning Image to Document Summary Attention for Zero-Shot Image Classification33
A Memory-Assisted Knowledge Transferring Framework with Curriculum Anticipation for Weakly Supervised Online Activity Detection32
Weighted Joint Distribution Optimal Transport Based Domain Adaptation for Cross-Scenario Face Anti-Spoofing32
Beyond Image Prior: Embedding Noise Prior into Latent Space of Conditional Denoising Transformer32
Predictive Display for Teleoperation Based on Vector Fields Using Lidar-Camera Fusion32
PartCom: Part Composition Learning for 3D Open-Set Recognition32
Shuffled Linear Regression with Outliers in Both Covariates and Responses32
Learning Box Regression and Mask Segmentation Under Long-Tailed Distribution with Gradient Transfusing32
SHARP: Shape-Aware Reconstruction of People in Loose Clothing32
Investigating Self-Supervised Methods for Label-Efficient Learning32
A Region-Based Randers Geodesic Approach for Image Segmentation31
Point-In-Context: Understanding Point Cloud via In-Context Learning31
InstaBoost++: Visual Coherence Principles for Unified 2D/3D Instance Level Data Augmentation31
Blur Invariants for Image Recognition31
LEO: Generative Latent Image Animator for Human Video Synthesis30
Object-Scene-Camera Decomposition and Recomposition for Data Efficient Monocular 3D Object Detection30
Uncertainty-Aware and Decoupled Distillation for Semantic Segmentation30
WildIng: A Wildlife Image Invariant Representation Model for Geographical Domain Shift30
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion30
Active Perception for Visual-Language Navigation30
Universal Facial Landmark Detection by Landmark-Clustering Relation-Reasoning Transformer29
Exemplar-Free Lifelong Person Re-identification via Prompt-Guided Adaptive Knowledge Consolidation29
TokenPacker: Efficient Visual Projector for Multimodal LLM29
CDistNet: Perceiving Multi-domain Character Distance for Robust Text Recognition29
Correction: Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement29
Deep Richardson–Lucy Deconvolution for Low-Light Image Deblurring29
An Optimal Transport View of Class-Imbalanced Visual Recognition29
Ultra-Lightweight Adaptive Bitrate Deep Video Compression28
A Family of Approaches for Full 3D Reconstruction of Objects with Complex Surface Reflectance28
Few-Shot Referring Video Single- and Multi-Object Segmentation Via Cross-Modal Affinity with Instance Sequence Matching28
Transformer-Based Context Condensation for Boosting Feature Pyramids in Object Detection28
Source-Free Domain Adaptation via Target Prediction Distribution Searching27
Day2Dark: Pseudo-Supervised Activity Recognition Beyond Silent Daylight27
Guest Editorial: Special Issue on Visual Datasets27
Reconstructing a Sphere and the Camera Focal Length from a Single View by Fitting Planes27
IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance27
SMPL-IKS: A Mixed Analytical-Neural Inverse Kinematics Solver for 3D Human Mesh Recovery27
Subspace Training Mitigates Gradient Noise Vulnerability27
Few-Shot Learning with Complex-Valued Neural Networks and Dependable Learning27
WildCLIP: Scene and Animal Attribute Retrieval from Camera Trap Data with Domain-Adapted Vision-Language Models27
HACG: Leveraging Hierarchical Alignment and Caption Generation for Text-Video Retrieval27
Editor’s Note: Special Issue on Computer Vision Approach for Animal Tracking and Modeling27
Evidence Conflict Sampling for Open-set Active Learning27
Editor’s Note: Special Issue on ACCV 202426
Polynomial Implicit Neural Framework for Promoting Shape Awareness in Generative Models26
Robust Image Restoration with an Adaptive Huber Function Based Fidelity26
Uniformity Preserving Transfer for Visual Prompt Tuning under Long-tailed Distribution26
Neural Architecture Search for Dense Prediction Tasks in Computer Vision26
Correction: Continual Face Forgery Detection via Historical Distribution Preserving26
An Interactive Conversational 3D Virtual Human26
A Novel Dataset and Lightweight Distillation Baseline for Highlight Transparent Object Detection26
Anti-Bandit for Neural Architecture Search26
AgMTR: Agent Mining Transformer for Few-Shot Segmentation in Remote Sensing25
GenderBias-VL: Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing25
Out-of-Distribution Detection with Virtual Outlier Smoothing25
Supervised Neural Style Transfer as an Augmentation Technique for Facial Landmark Detection25
Knowledge Distillation Meets Open-Set Semi-supervised Learning25
Geometric Prior Guided Feature Representation Learning for Long-Tailed Classification25
Hard-Normal Example-Aware Template Mutual Matching for Industrial Anomaly Detection25
Physics-Driven Spectrum-Consistent Federated Learning for Palmprint Verification24
Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding24
A Deeper Analysis of Volumetric Relightable Faces24
Zero-Shot Learning on 3D Point Cloud Objects and Beyond24
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning24
Vector-Symbolic Architecture for Event-Based Optical Flow24
CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer24
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary24
Generalized Relative Pose and Scale from Affine Correspondences24
On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook24
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering24
Relation-Guided Adversarial Learning for Data-Free Knowledge Transfer24
Preface to the Special Issue on Pattern Recognition (DAGM GCPR 2021)24
Unifying Viewgraph Sparsification and Disambiguation of Repeated Structures in Structure-from-Motion24
LLMFormer: Large Language Model for Open-Vocabulary Semantic Segmentation24
CamoVid60K: A Large-Scale Video Dataset for Moving Camouflaged Animals Understanding23
Multi-adversarial Faster-RCNN with Paradigm Teacher for Unrestricted Object Detection23
DustNet++: Deep Learning-Based Visual Regression for Dust Density Estimation23
Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast23
Correction: Variational Rectification Inference for Learning with Noisy Labels23
Structure-from-motion in micro-image domain for uncalibrated plenoptic 2.0 cameras23
FourierMIL: Fourier Filtering-based Multiple Instance Learning for Whole Slide Image Analysis23
Dynamic MAsk-Pruning Strategy for Source-Free Model Intellectual Property Protection23
SLNMapping: Super Lightweight Neural Mapping in Large-Scale Scenes23
Single-View View Synthesis with Self-rectified Pseudo-Stereo22
AnyPattern: Towards In-context Image Copy Detection22
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts22
Mining Generalized Multi-timescale Inconsistency for Detecting Deepfake Videos22
Sentimental Visual Captioning using Multimodal Transformer22
General Class-Balanced Multicentric Dynamic Prototype Pseudo-Labeling for Source-Free Domain Adaptation22
Defending Against Adversarial Examples Via Modeling Adversarial Noise22
Adversarial Learning Domain-Invariant Conditional Features for Robust Face Anti-spoofing22
Image-Based Virtual Try-On: A Survey22
UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection22
LiDAR-guided Geometric Pretraining for Vision-Centric 3D Object Detection22
Relative Norm Alignment for Tackling Domain Shift in Deep Multi-modal Classification22
Towards Generalized UAV Object Detection: A Novel Perspective from Frequency Domain Disentanglement21
A Comprehensive Study of the Robustness for LiDAR-Based 3D Object Detectors Against Adversarial Attacks21
Rethinking Open-World DeepFake Attribution with Multi-perspective Sensory Learning21
Semantically-aware Neural Radiance Fields for Visual Scene Understanding: A Comprehensive Review21
Visual-Geometric Collaborative Guidance for Affordance Learning21
Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline21
Correction: Scene Prior Filtering for Depth Super-Resolution21
Learning General and Specific Embedding with Transformer for Few-Shot Object Detection21
ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection21
RepSNet: A Nucleus Instance Segmentation Model Based on Boundary Regression and Structural Re-Parameterization20
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning20
Rethinking Out-of-Distribution Detection From a Human-Centric Perspective20
Towards Scene-Aware Video-to-Spatial Audio Generation20
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization20
A Generative Victim Model for Segmentation20
Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy20
Rethinking Neuromorphic Object Detection with Hybrid Dynamic Interaction Transformers20
Perspective-1-Ellipsoid: Formulation, Analysis and Solutions of the Camera Pose Estimation Problem from One Ellipse-Ellipsoid Correspondence20
CompViT: Real-Time Compressed Video Action Recognition with Asymmetric Transformer Networks20
Correction: Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding20
Not All Pixels are Equal: Learning Pixel Hardness for Semantic Segmentation20
Segment Anything in 3D with Radiance Fields20
IMC-Det: Intra–Inter Modality Contrastive Learning for Video Object Detection20
Thread Counting in Plain Weave for Old Paintings Using Regression Deep Learning Models20
Vision-Language Efficient Tuning for Mitigating Catastrophic Forgetting in Multi-Modal Learning20
Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization20
Evidential Robust Feature Learning for Generalized Few-Shot Segmentation19
A Survey of Multimodal Hallucination Evaluation and Detection19
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting19
Task Bias in Contrastive Vision-Language Models19
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation19
BayesAdapter: Enhanced Uncertainty Estimation in CLIP Few-Shot Adaptation19
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports19
0.18904805183411