Image and Vision Computing

Papers
(The TQCC of Image and Vision Computing is 9. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
ADVC: Adversarial dense video captioning with unsupervised pretraining462
Alignment and fusion for adaptive domain nighttime semantic segmentation207
Few-shot-based video generation via multimodal fusion and Fourier Spliter198
Feature decoupling and interaction network for defending against adversarial examples129
Modeling content-attribute preference for personalized image esthetics assessment128
GLMambaNet: Mamba-based decoder with local detail enhancement for semantic segmentation of remote sensing imagery107
Efficient ultra-lightweight convolutional attention network for embedded identity document recognition system97
Accurate and efficient salient object detection via position prior attention94
Multi-information guided camouflaged object detection87
G-TRACE: Grouped temporal recalibration for video object segmentation80
BF3D: Bi-directional fusion 3D detector with semantic sampling and geometric mapping78
Learning diverse and deep clues for person reidentification67
Lightweight multi-scale global attention enhancement network for image super-resolution67
DMNet: Image dehazing via Dual-Domain Modulation64
Active domain adaptation for semantic segmentation via dynamically balancing domainness and uncertainty64
Window normalization: Enhancing point cloud understanding by unifying inconsistent point densities62
Background debiased class incremental learning for video action recognition61
RGB-T tracking by modality difference reduction and feature re-selection59
AI-powered trustable and explainable fall detection system using transfer learning59
ABC: Aligning binary centers for single-stage monocular 3D object detection53
Hourglass cascaded recurrent stereo matching network52
DeepArUco++: Improved detection of square fiducial markers in challenging lighting conditions48
GAN-BodyPose: Real-time 3D human body pose data key point detection and quality assessment assisted by generative adversarial network48
HPD-Depth: High performance decoding network for self-supervised monocular depth estimation46
UCPNet: An Ultra-Lightweight Cross-Perception Network for Real-Time Semantic Segmentation44
Privacy-preserving explainable AI enable federated learning-based denoising fingerprint recognition model43
Single stage architecture for improved accuracy real-time object detection on mobile devices42
CODNet: Context-based object detection network for multimodal image captioning and virtual question answering42
MAFUNet: Mamba with adaptive fusion UNet for medical image segmentation41
SRMA-KD: Structured relational multi-scale attention knowledge distillation for effective lightweight cardiac image segmentation39
PST-Mamba: Spatio-temporal selective state fusion for effective point cloud video understanding with state space models39
Few-shot classification with multisemantic information fusion network38
Synthetic lidar point cloud generation using deep generative models for improved driving scene object recognition38
CAGS: Open-vocabulary 3D scene understanding with context-aware Gaussian splatting38
Recent advances in deterministic human motion prediction: A review36
SAGNet: Synergistic Attention-Graph Network For video salient object detection36
Two-stream transformer tracking with messengers36
Burst image super-resolution via multi-cross attention encoding and multi-scan state-space decoding36
A Point-2s reinforcement learning biomimetic model for estimating and analyzing human 3D motion posture34
SAFENet: Semantic-Aware Feature Enhancement Network for unsupervised cross-domain road scene segmentation34
1D kernel distillation network for efficient image super-resolution34
Deep learning with adaptive convolutions for classification of retinal diseases via optical coherence tomography34
Dual subspace clustering for spectral-spatial hyperspectral image clustering34
Editorial Board34
Editorial to special issue on selected extended works from 9th international conference on computer vision & image processing (CVIP) 202433
Utilizing Inherent Bias for Memory Efficient Continual Learning: A Simple and Robust Baseline33
Depth assisted novel view synthesis using few images32
FSBI: Deepfake detection with frequency enhanced self-blended images31
CSG-DOF:A Class Structure-Guided Discriminative Optimization Framework for few-shot object detection31
Frequency and content dual stream network for image dehazing31
CMS-net: Edge-aware multimodal MRI feature fusion for brain tumor segmentation31
RECALL+: Adversarial web-based replay for continual learning in semantic segmentation31
Memory-MambaNav: Enhancing object-goal navigation through integration of spatial–temporal scanning with state space models30
MVPCC-Net: Multi-View Based Point Cloud Completion Network for MLS data29
Visionary vigilance: Optimized YOLOV8 for fallen person detection with large-scale benchmark dataset29
ST-VTON: Self-supervised vision transformer for image-based virtual try-on29
Enhanced residual network for burst image super-resolution using simple base frame guidance28
Learning accurate monocular 3D voxel representation via bilateral voxel transformer28
Self-supervised Vision Transformers for 3D pose estimation of novel objects27
A spatial-frequency domain multi-branch decoder method for real-time semantic segmentation26
CNN and Transformer-based deep learning models for automated white blood cell detection26
TransMix: Crafting highly transferable adversarial examples to evade face recognition models26
Face deidentification with controllable privacy protection26
Editorial Board26
A multi-branch dual attention segmentation network for epiphyte drone images26
EMA-GS: Improving sparse point cloud rendering with EMA gradient and anchor upsampling25
SADGFeat: Learning local features with layer spatial attention and domain generalization25
RFSC-net: Re-parameterization forward semantic compensation network in low-light environments25
Editorial Board25
Deep learning enhanced monocular visual odometry: Advancements in fusion mechanisms and training strategies24
Mitigating human fall injuries: A novel system utilizing 3D 4-stream convolutional neural networks and image fusion24
Mixup Mask Adaptation: Bridging the gap between input saliency and representations via attention mechanism in feature mixup24
Landmark-in-facial-component: Towards occlusion-robust facial landmark localization24
Distributed collaborative machine learning in real-world application scenario: A white blood cell subtypes classification case study24
DynaGuide: A generalizable dynamic guidance framework for zero-shot guided unsupervised semantic segmentation24
Robust visual tracking via modified Harris hawks optimization24
UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification23
PatchMixer: Rethinking network design to boost generalization for 3D point cloud understanding23
Unsupervised Object Localization driven by self-supervised foundation models: A comprehensive review23
Feature alignment via mutual mapping for few-shot fine-grained visual classification23
Intelligent facial expression recognition and classification using optimal deep transfer learning model22
PAGML: Precise Alignment Guided Metric Learning for sketch-based 3D shape retrieval22
Doctor-in-the-Loop: An explainable, multi-view deep learning framework for predicting pathological response in non-small cell lung cancer22
Enhancing consistency in virtual try-on: A novel diffusion-based approach22
AGSAM-Net: UAV route planning and visual guidance model for bridge surface defect detection22
A survey on dynamic neural networks: From computer vision to multi-modal sensor fusion22
NPVForensics: Learning VA correlations in non-critical phoneme–viseme regions for deepfake detection22
Matte anything: Interactive natural image matting with segment anything model21
Underwater image restoration based on light attenuation prior and color-contrast adaptive correction21
Enhancing brain tumor classification in MRI images: A deep learning-based approach for accurate diagnosis21
A new multi-picture architecture for learned video deinterlacing and demosaicing with parallel deformable convolution and self-attention blocks21
Contrast enhancement of region of interest of backlit image for surveillance systems based on multi-illumination fusion21
Detection of anomaly in surveillance videos using quantum convolutional neural networks21
Dual-branch adaptive attention transformer for occluded person re-identification21
Object tracking based on temporal and spatial context information20
CollaborativeBEV: Collaborative bird eye view for reconstructing crowded environment20
Phase shift guided dynamic view synthesis from monocular video20
OFACD: An end-to-end change detection network for small UAVs remote sensing with viewpoint differences20
STAFFormer: Spatio-temporal adaptive fusion transformer for efficient 3D human pose estimation20
Deep learning-based efficient diagnosis of periapical diseases with dental X-rays20
Self-knowledge distillation based on knowledge transfer from soft to hard examples20
Anchor-based discriminative dual distribution calibration for transductive zero-shot learning20
SDMNet: Spatially dilated multi-scale network for object detection for drone aerial imagery20
E-Net for pansharpening: A super-resolution perspective20
M2VAD: Multiview multi20
Multi-axis interactive multidimensional attention network for vehicle re-identification19
WPE: Weighted prototype estimation for few-shot learning19
PW-NeRF: Progressive wavelet-mask guided neural radiance fields view synthesis19
TQRFormer: Tubelet query recollection transformer for action detection19
Distribution-modulated binary neural network for image classification19
A novel facial expression recognition model based on harnessing complementary features in multi-scale network with attention fusion19
PixTention: Dynamic pixel-level adapter using attention maps19
CSUnetr: Cross-scale attention based U-Net transformers for whole-brain segmentation with targeted hippocampal analysis in brain MR images19
Video anomaly detection based on a multi-layer reconstruction autoencoder with a variance attention strategy18
Attentive spatial-temporal contrastive learning for self-supervised video representation18
SAMNet: Adapting segment anything model for accurate light field salient object detection18
Adaptive graph reasoning network for object detection18
Estimating blood pressure using video-based PPG and deep learning18
FgbCNN: A unified bilinear architecture for learning a fine-grained feature representation in facial expression recognition18
Adaptive scale matching for remote sensing object detection based on aerial images18
Source domain prior-assisted segment anything model for single domain generalization in medical image segmentation18
Class-discriminative domain generalization for semantic segmentation18
Real-time gait biometrics for surveillance applications: A review18
RLTNT: An explainable residual learning-based transformer model for kidney disease classification18
AHA-track: Aggregating hierarchical awareness features for single18
Real-time human-centric segmentation for complex video scenes18
ECNet: An edge-guided and cross-image perception network for collaborative camouflaged object detection17
Online multi-object tracking with δ-GLMB filter based on occlusion and identity switch handling17
Corrigendum to “A novel framework for diverse video generation from a single video using frame-conditioned denoising diffusion probabilistic model and ConvNeXt-V2” [Image and Vision Computing 154 (20217
Unleashing spatial-awareness for robust object tracking17
Optimal deep transfer learning based ethnicity recognition on face images17
H-net: Unsupervised domain adaptation person re-identification network based on hierarchy17
Face and body-shape integration model for cloth-changing person re-identification17
Social robot in service of the cognitive therapy of elderly people: Exploring robot acceptance in a real-world scenario17
Semantic-aware for point cloud domain adaptation with self-distillation learning17
DFG-HCEN: A distinctive-feature guided and hierarchical channel enhanced network-based infrared and visible image fusion17
Dynamic semantic prototype perception for text–video retrieval17
Data-driven 2D-EWT based diabetic retinopathy identification using hybrid neural network17
Enhancing small object tracking with reversible rescaling networks17
Your image generator is your new private dataset17
PR-DETR: Extracting and utilizing prior knowledge for improved end-to-end object detection17
CRFormer: A cross-region transformer for shadow removal17
Stacked graph bone region U-net with bone representation for hand pose estimation and semi-supervised training17
TABNet: A Triplet Augmentation Self-recovery framework with Boundary-aware Pseudo-labels for scribble-based medical image segmentation17
Few-shot class incremental learning via prompt transfer and knowledge distillation17
Editorial Board17
Dual-stage network combining transformer and hybrid convolutions for stereo image super-resolution17
An edge-aware high-resolution framework for camouflaged object detection16
Perceiving local relative motion and global correlations for weakly supervised group activity recognition16
BCDPose: Diffusion-based 3D Human Pose Estimation with bone-chain prior knowledge16
Point-cloud-based hand gesture recognition using principal component analysis and boundary extraction16
A lightweight shallow convolution neural network for automatic identification of Diabetic Foot Ulcers16
Exploiting spatial and temporal context for online tracking with improved transformer15
CVAD-GAN: Constrained video anomaly detection via generative adversarial network15
Similarity verification of kinship pairs using metricized emphasis15
Speaker independent VSR: A systematic review and futuristic applications15
Semantic scene graph generation based on an edge dual scene graph and message passing neural network15
Synthetic multi-view clustering with missing relationships and instances15
A comprehensive survey on magnetic resonance image reconstruction15
Cross-level fusion network for two-stage polyp segmentation via integrity learning15
UIR-ES: An unsupervised underwater image restoration framework with equivariance and stein unbiased risk estimator15
Editorial Board15
Self-distillation guided Semantic Knowledge Feedback network for infrared–visible image fusion15
SDE-RAE:CLIP-based realistic image reconstruction and editing network using stochastic differential diffusion15
Bridging efficiency and interpretability: Explainable AI for multi-classification of pulmonary diseases utilizing modified lightweight CNNs15
HMPFormer: Hierarchical vision transformer with multi-perspective feature learning for precise polyp segmentation15
Resource-aware strategies for real-time multi-person pose estimation15
Guest Editorial : Learning with Manifolds in Computer Vision15
AI4RDD: Artificial Intelligence and Rare Disease Diagnosis: A proposal to improve the anamnesis process14
Synthetic data sets for person Re-Identification: A critical analysis14
Enhancing UAV small target detection: A balanced accuracy-efficiency algorithm with tiered feature focus14
BTMTrack: Robust RGB-T tracking via dual-template bridging and temporal-modal candidate elimination14
Optimizing multimodal personalized disease prediction accuracy using generated prompts and large language models14
Deep hybrid learning for facial expression binary classifications and predictions14
PD-DDPM: Prior-driven diffusion model for single image dehazing14
Efficient Mamba: Overcoming the visual limitations of Mamba with innovative structures14
Twin relaxed least squares regression with classwise mean constraint for image classification14
Preserving instance-level characteristics for multi-instance generation14
Editorial Board14
External knowledge-assisted Transformer for image captioning14
Weather-degraded image semantic segmentation with multi-task knowledge distillation14
CoHAtNet: An integrated convolutional-transformer architecture with hybrid self-attention for end-to-end camera localization14
Flexible multi-objective particle swarm optimization clustering with game theory to address human activity discovery fully unsupervised14
DMC-former: A dual-flow dynamic mask and collaborative attention-based network for micro-expression recognition14
Black-box reversible adversarial examples with invertible neural network13
GFFT: Global-local feature fusion transformers for facial expression recognition in the wild13
DRM-YOLO: A YOLOv11-based structural optimization method for small object detection in UAV aerial imagery13
Fuzzy set-based Bernoulli Random Noise Weighted Loss for unsupervised person re-identification13
Feature extraction and fusion algorithm for infrared visible light images based on residual and generative adversarial network13
Contrastive learning based facial action unit detection in children with hearing impairment for a socially assistive robot platform13
Parameter efficient finetuning of text-to-image models with trainable self-attention layer13
Qualitative failures of image generation models and their application in detecting deepfakes13
Multi-object tracking with adaptive measurement noise and information fusion13
Editorial Board13
Knowledge graph construction in hyperbolic space for automatic image annotation13
Video object segmentation by multi-scale attention using bidirectional strategy13
Editorial Board13
A dual-channel network based on occlusion feature compensation for human pose estimation13
A decision support system for acute lymphoblastic leukemia detection based on explainable artificial intelligence13
A video anomaly detection and classification method based on cross-modal feature alignment13
A dedicated benchmark for contour-based corner detection evaluation13
Transformer-based feature interactor for person re-identification with margin self-punishment loss13
Cross-modal hybrid architectures for gastrointestinal tract image analysis: A systematic review and futuristic applications13
DiPS: Discriminative pseudo-label sampling with self-supervised transformers for weakly supervised object localization13
Gait recognition via View-aware Part-wise Attention and Multi-scale Dilated Temporal Extractor12
Boosting semi-supervised face recognition with raw faces12
Bidirectional causal learning for visual question answering12
Robust ensemble person reidentification via orthogonal fusion with occlusion handling12
SEAD-Net:Complex underwater image segmentation via semantic-enhanced and detail-aware collaboration12
ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentation12
Action-aware anchor-based frame selection strategy for action recognition12
Unified Volumetric Avatar: Enabling flexible editing and rendering of neural human representations12
Efficient masked feature and group attention network for stereo image super-resolution12
Depth awakens: A depth-perceptual attention fusion network for RGB-D camouflaged object detection12
A gan inversion-based multimodal framework for micro-expression recognition12
Effective hybrid attention network based on pseudo-color enhancement in ultrasound image segmentation12
Editorial Board12
Drone-NeRF: Efficient NeRF based 3D scene reconstruction for large-scale drone survey12
OCUCFormer: An Over-Complete Under-Complete Transformer Network for accelerated MRI reconstruction12
ECT: Fine-grained edge detection with learned cause tokens12
RGB road scene material segmentation12
Tri-UNetX: Tri-plane UNet with xLSTM for 3D cell segmentation12
Dense open-set recognition based on training with noisy negative images11
Improving defocus blur detection via adaptive supervision prior-tokens11
Universal domain adaptation from multiple black-box sources11
Hierarchical spatiotemporal Feature Interaction Network for video saliency prediction11
A supervised approach for the detection of AM-FM signals’ interference regions in spectrogram images11
CF-SOLT: Real-time and accurate traffic accident detection using correlation filter-based tracking11
Multi-level feature disentanglement network for cross-dataset face forgery detection11
Leveraging spatial-channel attention in U-Net for enhanced segmentation of martian dust storms11
Image–text feature learning for unsupervised visible–infrared person re-identification11
On the relevance of patch-based extraction methods for monocular depth estimation11
STIFormer: RGB-T tracking via Spatial–Temporal Interaction Transformer11
MOT-STM: Maritime Object Tracking: A Spatial-Temporal and Metadata-based approach11
LELD: Learn enhancement by learning degradation11
Text-augmented Multi-Modality contrastive learning for unsupervised visible-infrared person re-identification11
Modal-aware contrastive learning for hyperspectral and LiDAR classification11
EFDCNet: Encoding fusion and decoding correction network for RGB-D indoor semantic segmentation11
Rethinking the sample relations for few-shot classification11
Federated learning based nonlinear two-stage framework for full-reference image quality assessment: An application for biometric11
SAKD: Sparse attention knowledge distillation11
A lightweight hash-directed global perception and self-calibrated multiscale fusion network for image super-resolution11
GW-net: An efficient grad-CAM consistency neural network with weakening of random erasing features for semi-supervised person re-identification11
Three dimensional tracking of rigid objects in motion using 2D optical flows11
SAMUNet: Enhancing pillar-based 3D object detection in autonomous driving with Shape-aware Mini-Unet11
Distributed quantum model learning for traffic density estimation11
Attention guided multi-level feature aggregation network for camouflaged object detection11
AES-Net: An adapter and enhanced self-attention guided network for multi-stage glaucoma classification using fundus images11
Editorial Board11
Machine learning applications in breast cancer prediction using mammography11
0.10170221328735