Computer Vision and Image Understanding

Papers
(The TQCC of Computer Vision and Image Understanding is 7. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Luminance prior guided Low-Light 4C catenary image enhancement410
Editorial Board132
Editorial Board121
Improving the planarity and sharpness of monocularly estimated depth images using the Phong reflection model118
Editorial Board107
Exploring using jigsaw puzzles for out-of-distribution detection90
Extending function mixture network for improved spectral super-resolution74
MATTE: Multi-task multi-scale attention67
Efficient cross-information fusion decoder for semantic segmentation56
Editorial Board56
3D semantic segmentation based on spatial-aware convolution and shape completion for augmented reality applications56
Lightweight feature point detection network with channel enhancement56
Editorial Board56
Emerging image generation with flexible control of perceived difficulty55
Modality adaptation via feature difference learning for depth human parsing53
QB-MOTR: A simple query bootstrapping end-to-end multi-object tracking method with transformer52
Siamese self-supervised learning for fine-grained visual classification47
Deducing health cues from biometric data45
SNRD-Net: SNR-aware dual enhancement network for low-light images45
RetSeg3D: Retention-based 3D semantic segmentation for autonomous driving45
Twin-SegNet: Dynamically coupled complementary segmentation networks for generalized medical image segmentation44
Spatial Sensitive Grad-CAM++: Towards High-Quality Visual Explanations for Object Detectors via Weighted Combination of Gradient Maps43
REST: A resolution preserving network for photorealistic style transfer via semantic distillation42
Vision-based mistake analysis in procedural activities: A review of advances and challenges40
CRML-Net: Cross-Modal Reasoning and Multi-Task Learning Network for tooth image segmentation40
JEMA: Joint Embedding of Multimodal and multi-view Alignment in human-centric embedding space for manufacturing39
NaviFormer: Multimodal scene segmentation for assistive navigation39
Exploring the differences in adversarial robustness between ViT- and CNN-based models using novel metrics38
Robust Teacher: Self-correcting pseudo-label-guided semi-supervised learning for object detection36
Feature reconstruction and metric based network for few-shot object detection36
Convolutional neural network framework for deepfake detection: A diffusion-based approach34
RelFormer: Advancing contextual relations for transformer-based dense captioning32
Feature preserving 3D mesh denoising with a Dense Local Graph Neural Network32
PConvSRGAN: Real-world super-resolution reconstruction with pure convolutional networks31
Syntactically and semantically enhanced captioning network via hybrid attention and POS tagging prompt30
Improved Short-term Dense Bottleneck network for efficient scene analysis28
A lightweight and robust framework for small object detection in UAV imagery28
Editorial Board27
Iterative Caption Generation with Heuristic Guidance for enhancing knowledge-based visual question answering27
CCNeXt: An effective self-supervised stereo depth estimation approach27
Embedding AI ethics into the design and use of computer vision technology for consumer’s behaviour understanding27
Editorial Board26
SIERRA: A robust bilateral feature upsampler for dense prediction26
View-aligned pixel-level feature aggregation for 3D shape classification26
Hierarchical contrastive distillation: Bridging multi-level semantics for enhanced knowledge transfer25
Adaptive CNN filter pruning using global importance metric25
Implicit and explicit commonsense for multi-sentence video captioning25
Towards efficient image and video style transfer via distillation and learnable feature transformation25
GaitBranch: A multi-branch refinement model combined with frame-channel attention mechanism for gait recognition25
Prompt-guided dual-path UNet with Mamba for medical image segmentation24
Are Candidate Models Really Needed for Active Learning?24
SDC-Net: A novel selective dilated convolution network for medical images segmentation24
Redundancy-aware memory update for improved video object segmentation24
Reverse Stable Diffusion: What prompt was used to generate this image?24
3D object feature extraction and classification using 3D MF-DFA23
M3A: A multimodal misinformation dataset for media authenticity analysis22
Attribute-guided Relevance Propagation for interpreting image classifier based on Deep Neural Networks22
Other tokens matter: Exploring global and local features of Vision Transformers for Object Re-Identification22
Pseudo initialization based Few-Shot Class Incremental Learning22
Hi-ROS: Open-source multi-camera sensor fusion for real-time people tracking22
M-Control: Improving text–image consistency via Mask-Guided ControlNet22
Lightning fast video anomaly detection via multi-scale adversarial distillation22
Editorial Board21
Lv-Adapter: Adapting Vision Transformers for Visual Classification with Linear-layers and Vectors21
Learning spectral transform for 3D human motion prediction21
Editorial Board21
LARKED:A lightweight and reliable keypoint detection method for feature matching20
Unsupervised real image super-resolution via knowledge distillation network20
Dynamic deep multi-label image data augmentation based on self-paced learning20
Extensions in channel and class dimensions for attention-based knowledge distillation20
BARD: A Basketball Action Recognition Dataset for multi-label classification20
Enhanced dual contrast representation learning with cell separation and merging for breast cancer diagnosis20
TFUT: Task fusion upward transformer model for multi-task learning on dense prediction20
Self-supervised network for low-light traffic image enhancement based on deep noise and artifacts removal19
An efficient direct solution of the perspective-three-point problem19
Continuous fake media detection: Adapting deepfake detectors to new generative techniques19
Enhancing feature representation in Siamese networks for object tracking with ranking-based loss19
Uncertainty estimation using boundary prediction for medical image super-resolution19
A multi-view-CNN framework for deep representation learning in image classification18
Ensemble learning-based method for maritime background subtraction in open sea environments18
MOSAIC: Maximizing out-of-distribution sensitivity via aligned image classification18
Hexagonal mesh-based neural rendering for real-time rendering and fast reconstruction18
SSDA-YOLO: Semi-supervised domain adaptive YOLO for cross-domain object detection18
UniMultNet: Action recognition method based on multi-scale feature fusion and video-text constraint guidance18
A dynamic hybrid network with attention and mamba for image captioning18
When super-resolution meets camouflaged object detection: A comparison study18
Multi-dimensional attention-aided transposed ConvBiLSTM network for hyperspectral image super-resolution17
Multi-view cognition with path search for one-shot part labeling17
MOSAIC: A multi-view 2.5D organ slice selector with cross-attentional reasoning for anatomically-aware CT localization in medical organ segmentation16
Indoor UAV navigation using event cameras and intermediate frame reconstruction16
Few-shot Medical Image Segmentation via Boundary-extended Prototypes and Momentum Inference16
SPSC-Net: Shared parallel space-channel attention mechanism transformer network for cell sequence image segmentation16
Editorial Board16
Statistical-driven adaptive data augmentation for single-domain generalized object detection16
Casting a BAIT for offline and online source-free domain adaptation15
A robust kinship verification scheme using face age transformation15
BasicTAD: An astounding RGB-Only baseline for temporal action detection15
Deep parametric Retinex decomposition model for low-light image enhancement15
Global key knowledge distillation framework15
SHOWMe: Robust object-agnostic hand-object 3D reconstruction from RGB video15
CTM: Cross-time temporal module for fine-grained action recognition15
Sketch-based 3D shape retrieval via teacher–student learning15
Transformed ROIs for capturing visual transformations in videos15
Scribble-based complementary graph reasoning network for weakly supervised salient object detection14
Editorial Board14
MLGPnet: Multi-granularity neural network for 3D shape recognition using pyramid data14
Semantic-driven diffusion for sign language production with gloss-pose latent spaces alignment14
The shading isophotes: Model and methods for Lambertian planes and a point light14
3D Pose Nowcasting: Forecast the future to improve the present14
Real-time distributed video analytics for privacy-aware person search14
Attention-induced semantic and boundary interaction network for camouflaged object detection14
Biometric technology roadmapping for personalized augmentative and alternative communication14
XLITE-Unet: Extremely Light and Efficient Deep learning architecture with selective atrous and axial depthwise convolution for image segmentation13
For a semiotic AI: Bridging computer vision and visual semiotics for computational observation of large scale facial image archives13
Semi-supervised Cycle-GAN for face photo-sketch translation in the wild13
Towards robust 3D human reconstruction with uncertainty-aware low-rank adaptation13
Rethink arbitrary style transfer with transformer and contrastive learning13
Quantifying model uncertainty for semantic segmentation of Fluorine-19 MRI using stochastic gradient MCMC13
Multiscale Spatio-Temporal Fusion Network for video dehazing13
MASK_LOSS guided non-end-to-end image denoising network based on multi-attention module with bias rectified linear unit and absolute pooling unit13
Image denoising model based on edge structure guidance and frequency-domain/spatial-domain coupling13
Real-time fusion of stereo vision and hyperspectral imaging for objective decision support during surgery13
DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation13
Combinational sign language recognition13
Space–time recurrent memory network13
EADA: Efficient adaptive data augmentation13
Extending class activation mapping using Gaussian receptive field13
OFCA-Net: An explainable optical flow-based framework for face forgery detection13
A LLM-guided hybrid Mamba-Transformer architecture for part-to-whole motion synthesis13
Learning representational invariances for data-efficient action recognition13
High-speed autonomous flight and obstacle avoidance for quadrotors in unknown dynamic environments based on imitation learning13
Tensor robust PCA with nonconvex and nonlocal regularization12
To make yourself invisible with Adversarial Semantic Contours12
VIDF-Net: A Voxel-Image Dynamic Fusion method for 3D object detection12
Generalized prompt-driven zero-shot domain adaptive segmentation with feature rectification and semantic modulation12
A multi camera unsupervised domain adaptation pipeline for object detection in cultural sites through adversarial learning and self-training12
α-EGAN: α-Energy distance GAN with an early stopping rule12
Feature-aligned distillation for dense object detection via refined semantic guidance and distribution consistency12
Dual cross-enhancement network for highly accurate dichotomous image segmentation12
Edge-aware graph reasoning network for image manipulation localization12
Local Consistency Guidance: Personalized Stylization Method of Face Video12
Distributed multi-target tracking and active perception with mobile camera networks12
Semantic manipulation through the lens of Geometric Algebra12
Robust attention ranking architecture with frequency-domain transform to defend against adversarial samples12
HFINet: Hybrid Feature Integration for enhancing collaborative camouflaged object detection12
FAR-AMTN: Attention Multi-Task Network for Face Attribute Recognition12
LightSOD: Towards lightweight and efficient network for salient object detection11
Exploring black-box adversarial attacks on Interpretable Deep Learning Systems11
Semantically accurate super-resolution Generative Adversarial Networks11
Bi-granularity balance learning for long-tailed image classification11
Constructing adaptive spatial-frequency interactive network with bi-directional adapter for generalizable face forgery detection11
FAM: Improving columnar vision transformer with feature attention mechanism11
4DHumanOutfit: A multi-subject 4D dataset of human motion sequences in varying outfits exhibiting large displacements11
GAN inversion via cross-domain feature fusion and invertibility decomposition11
EPDiff: Enhancing Prior-guided Diffusion model for Real-world Image Super-Resolution11
BiPG-FER: Bi-intelligence probabilistic graph for facial expression inference drived by action units11
Object re-identification via spatial–temporal fusion networks and causal identity matching11
An effective CNN and Transformer fusion network for camouflaged object detection11
Self-supervision & meta-learning for one-shot unsupervised cross-domain detection11
Editorial Board11
Context perturbation: A Consistent alignment approach for Domain Adaptive Semantic Segmentation11
Discriminative object tracking by domain contrast11
Accurate depth image generation via overfit training of point cloud registration using local frame sets11
Looking into the unknown: Exploring Action Discovery for segmentation of known and unknown actions11
DM-Align: Leveraging the power of natural language instructions to make changes to images11
GSNNet: Group semantic-guided neighbor interaction network for co-salient object detection11
EFSCNN: Encoded Feature Sphere Convolution Neural Network for fast non-rigid 3D models classification and retrieval11
Minimum error adaptive RGB calibration in a context of colorimetric uncertainty for cultural heritage preservation11
LocoGAN — Locally convolutional GAN11
Deep learning-based estimation of whole-body kinematics from multi-view images11
Certifiable algorithms for the two-view planar triangulation problem11
Editorial Board11
Lifelong visible–infrared person re-identification via replay samples domain-modality-mix reconstruction and cross-domain cognitive network11
Survey on fast dense video segmentation techniques11
GradPaint: Gradient-guided inpainting with diffusion models11
UATST: Towards unpaired arbitrary text-guided style transfer with cross-space modulation11
Comprehensive regional guidance for attention map semantics in text-to-image diffusion models11
UGLF: Uncertainty-Gated Global–Local Fusion for generalizable face forgery detection11
Learning rotation equivalent scene representation from instance-level semantics: A novel top-down perspective11
MFCT: Multi-Frequency Cascade Transformers for no-reference SR-IQA10
Bidirectional brain image translation using transfer learning from generic pre-trained models10
An efficient three-stage network via Multi-Scale Orthogonal Complementary Transformer for low-light image enhancement10
Underwater image quality evaluation via deep meta-learning: Dataset and objective method10
AWADA: Foreground-focused adversarial learning for cross-domain object detection10
Exploring joint embedding predictive architectures for pretraining convolutional neural networks10
Periocular biometrics and its relevance to partially masked faces: A survey10
Adaptive semantic guidance network for video captioning10
Joint coupled dictionaries-based visible-infrared image fusion method via texture preservation structure in sparse domain10
Evaluating the effect of image quantity on Gaussian Splatting: A statistical perspective10
Real-world efficient fall detection: Balancing performance and complexity with FDGA workflow10
On the coherency of quantitative evaluation of visual explanations10
Phase-based video motion magnification with handheld cameras10
Editorial Board10
Editorial Board10
View consistency aware holistic triangulation for 3D human pose estimation10
Adaptive gradients and weight projection based on quantized neural networks for efficient image classification10
MDC-Net: Multi-domain constrained kernel estimation network for blind image super resolution10
Adversarial Style Mixup and Improved Temporal Alignment for Cross-Domain Few-Shot Action Recognition10
EAUAV-YOLO: An efficiency and accuracy enhanced lightweight detector for small objects in aerial images10
MS2MUnet: A Multi-Scale Spectral Mamba U-Net for medical image segmentation10
S2DNet: A self-supervised deraining network using monocular videos10
OVGrasp: Open-Vocabulary Intent Detection for Grasping Assistance using ExoGlove9
Disentangled generation network for enlarged license plate recognition and a unified dataset9
TEMSA:Text enhanced modal representation learning for multimodal sentiment analysis9
Lightweight cross-modal transformer for RGB-D salient object detection9
SASFNet: Soft-edge awareness and spatial-attention feedback deep network for blind image deblurring9
Adaptive feature denoising based deep convolutional network for single image super-resolution9
MKP-Net: Memory knowledge propagation network for point-supervised temporal action localization in livestreaming9
Made-In: An immersive human-in-the-loop analytics platform for enhancing creative processes in fashion9
A closer look at branch classifiers of multi-exit architectures9
Fourier analysis on robustness of graph convolutional neural networks for skeleton-based action recognition9
Human skeletons and change detection for efficient violence detection in surveillance videos9
Incorporating degradation estimation in light field spatial super-resolution9
NeRFtrinsic Four: An end-to-end trainable NeRF jointly optimizing diverse intrinsic and extrinsic camera parameters9
Self-supervised vision transformers for semantic segmentation9
An image denoising method based on the nonlinear Schrödinger equation and spectral subband decomposition9
Certifiable planar relative pose estimation with gravity prior9
Multimodal transformer–diffusion framework for large-scale reconstruction of soccer tracking data9
Multi-person 3D pose estimation from a single image captured by a fisheye camera9
AnomalySD: One-for-all few-shot anomaly detection via pre-trained diffusion models9
Once Upon a Goal: Towards orientation-based shot metrics in football9
CMGNet: Collaborative multi-modal graph network for video captioning9
Constituent Attention for Vision Transformers9
Editorial Board9
MAL-Net: Multiscale Attention Link Network for accurate eye center detection9
Editorial Board8
Text-Aided Domain Adaptation for CLIP-like models and application to challenging domain shifts8
Editorial Board8
A real-time image super-resolution model based on U-shaped deep feature extraction module8
Time-archival camera virtualization for sports and visual performances8
FDPAdapter : Adapting segment anything in challenging vision tasks via frequency-domain priors8
Modality mixer exploiting complementary information for multi-modal action recognition8
Progressive multi-scale fusion network for RGB-D salient object detection8
: Localized text prompt refinement for zero-shot referring image segmentation8
Channel-aware feature mining network for Visible–Infrared Person Re-identification8
Cascading attention enhancement network for RGB-D indoor scene segmentation8
Learning key lines for multi-object tracking8
A survey on class-agnostic counting: Advancements from reference-based to open-world text-guided approaches8
Conditioning diffusion models via attributes and semantic masks for face generation8
Dual adversarial model: Exploring low-dimensional space features for point clouds generating and completing8
2.5D visual relationship detection8
Bypass network for semantics driven image paragraph captioning8
Blur aware metric depth estimation with multi-focus plenoptic cameras8
Distribution-aware contrastive learning for domain adaptation in 3D LiDAR segmentation8
Continual learning on 3D point clouds with random compressed rehearsal8
Sparse graph matching network for temporal language localization in videos8
EARS4SEE: A multimodal audio description system dedicated to blind and visually impaired users8
Invisible backdoor attack with attention and steganography8
Multimodal vs. unimodal approaches to uncertainty in 3D image segmentation under distribution shifts8
Leaf cultivar identification via prototype-enhanced learning8
0.095232963562012