Machine Vision and Applications

Papers
(The median citation count of Machine Vision and Applications is 2. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
StyleDemorpher: high-quality face demorphing via StyleGAN2’s latent space85
Class-aware cross-domain target detection based on cityscape in fog85
Development of a robust cascaded architecture for intelligent robot grasping using limited labelled data61
Non-contact SpO2 monitoring via multi-channel pulse signals from facial videos using machine learning51
ECM: arbitrary style transfer via Enhanced-Channel Module38
A method for high dynamic range 3D color modeling of objects through a color camera25
DMU-Net: a dual stream multi-scale U-Net for image splicing forgery localization24
An integration of deep network with random forests framework for image quality assessment in real-time23
Medtransnet: advanced gating transformer network for medical image classification22
Text-driven object affordance for guiding grasp-type recognition in multimodal robot teaching22
SBAHGNet:3D human pose estimation via skeleton-biased attention and high-frequency enhanced graph convolution21
Amodal segmentation for occlusion-aware instance recovery of juvenile abalone21
Lightweight image dehazing via physics-guided neural networks21
End-to-end unsupervised learning of latent-space clustering for image segmentation via fully dense-UNet and fuzzy C-means loss20
Real estate pricing prediction via textual and visual features18
Automated flow-mediated dilation using neural network-based segmentation17
Obs-tackle: an obstacle detection system to assist navigation of visually impaired using smartphones17
A hybrid overlapping group sparsity denoising model with fractional-order total variation and non-convex regularizer17
Innovative surface roughness detection method based on white light interference images16
Using breast density for hybrid region and pixel-level loss function16
Enforced clustering for zero-to-one-shot texture anomaly detection16
Editing implicit and explicit representations of radiance fields: a survey15
Enhancing hyperspectral image classification: DeepXTE for efficient semantic feature extraction15
A stereo vision SLAM with moving vehicles tracking in outdoor environment15
MSPKD: multi spatial projectors for knowledge distillation in semantic segmentation15
Motion-region annotation for complex videos via label propagation across occluders14
DPA-Net: a real-time dual-pyramid attention network for UAV object detection14
A motion direction detecting model for colored images based on the Hassenstein–Reichardt model14
Tcdgnet: a texture and chromaticity dual-guided network for color document super-resolution14
L-VAE: variational auto-encoder with learnable beta for disentangled representation13
LOID: Lane Occlusion Inpainting and Detection for Enhanced Autonomous Driving Systems13
Decoupling spatial and spectral features for efficient hyperspectral image super-resolution13
Generalized few-shot learning under large scope by using episode-wise regularizing imprinting12
Ubiquitous vision of transformers for person re-identification12
CGA-Net: channel-wise gated attention network for improved super-resolution in remote sensing imagery12
Alternate guidance network for boundary-aware camouflaged object detection12
Axes-aligned non-linear optimized PnP algorithm12
AFC-Net: adjacent feature complementary for crowded pedestrian detection12
Generation of realistic synthetic cable images to train deep learning segmentation models12
Modeling driving task-relevant attention for intelligent vehicles using triplet ranking12
Discriminant distance template matching for image recognition12
Multi-feature fusion network based on wavelet transform and multi-scale cross-response for hyperspectral image classification12
A dual progressive strategy for long-tailed visual recognition11
LLM-augmented semantic reasoning for robust AUV docking under uncertain environments11
Correction: Unsupervised single-shot depth estimation using perceptual reconstruction11
Improving knowledge distillation via pseudo-multi-teacher network11
Novel Cauchy mixture modeling combined with the Sparse-RCNN architecture for enhanced multi-person pose estimation11
A multi-modal framework for continuous and isolated hand gesture recognition utilizing movement epenthesis detection11
Two-stage structural information enhancement for source-free domain adaptation11
Kernel based local matching network for video object segmentation11
Specular Surface Detection with Deep Static Specular Flow and Highlight11
A general two-stage framework of tensor low-rank representation for enhanced image denoising and clustering10
RPIM-net: residual channel prior-driven interaction multi-scale network for stereo image deraining10
CAMTrack: a combined appearance-motion method for multiple-object tracking10
Benchmarking large and small MLLMs10
SGL-SLAM: a semantic and geometric RGB-D visual SLAM enhanced with line features for dynamic environments10
Traversing the subspace of adversarial patches10
Twinned attention network for occlusion-aware facial expression recognition10
Redundancy-free label space and dual-feature collaboration for multi-label feature selection10
Camera-based mapping in search-and-rescue via flying and ground robot teams10
Enhanced hyperspectral image reconstruction via parallel 2D/3D convolution with global layer purification and multiscale pooling fusion10
Imstrack: infrared maritime small target tracking with adaptive scaling field view9
Visually-guided audio-visual aegmentation via multi-scale fusion and content-guided attention9
Deep learning for analysis of visible-light polarimetric image: a review9
Audio-visual localization based on spatial relative sound order9
Correction: Real estate pricing prediction via textual and visual features9
Triple-attention enhanced and RepViT-driven LiDAR 3D object detection for complex traffic scenarios9
Enhanced point cloud processing through geometric affine transformations and curvature-based sampling9
Generating comprehensive scene graphs with integrated multiple attribute detection9
GOA-net: generic occlusion aware networks for visual tracking9
Adversarial imitation learning-based network for category-level 6D object pose estimation9
Online continual learning with saliency-guided experience replay using tiny episodic memory9
3D face parsing based on 2D CPFNet: conformal parameterized face parsing network9
LDNet: low-light image enhancement with joint lighting and denoising9
Shape related unknown object one-shot learning grasping9
OmniGlasses: an optical aid for stereo vision CNNs to enable omnidirectional image processing9
FOCUS: Frequency-Optimized Conditioning of diffUSion models for mitigating catastrophic forgetting during test-time adaptation8
Explainable interactive projections of images8
IoU-aware feature fusion R-CNN for dense object detection8
Multi-scale convolution underwater image restoration network8
Shape description losses for medical image segmentation8
Chfnet: a coarse-to-fine hierarchical refinement model for monocular depth estimation8
Reflection removal using recurrent polarization-to-polarization network8
A comprehensive survey on SLAM and machine learning approaches for indoor autonomous navigation of mobile robots8
X-Align++: cross-modal cross-view alignment for Bird’s-eye-view segmentation8
Overcoming occlusions in AR, via multi-view, real-time 3D human pose estimation8
A camera style-invariant learning and channel interaction enhancement fusion network for visible-infrared person re-identification8
Fusing bilinear multi-channel gated vector for fine-grained classification8
MÆIDM: multi-scale anomaly embedding inpainting and discrimination for surface anomaly detection8
A multi-class segmentation algorithm for oral and maxillofacial structures in CBCT images: two-level attention mechanism and optimized loss8
Cross-dataset video deepfake detection using Transformer and CNN architectures7
Visual-inertial SLAM with line segment merging and efficient feature tracking method7
Enhancing object SLAM for outdoor environments: robust reconstruction and relocalization7
Pakistan sign language recognition: leveraging deep learning models with limited dataset7
An efficient ground segmentation approach for LiDAR point cloud utilizing adjacent grids7
Evolving brain tumor segmentation: differential evolution-optimized ensemble deep learning for multi-modal MRI analysis7
Evolution algorithm of parametric active contour model based on Gaussian smoothing filter7
Mobgazenet: robust gaze estimation mobile network based on progressive attention mechanisms7
Real-time pedestrian pose estimation, tracking and localization for social distancing7
An adaptive interpolation and 3D reconstruction algorithm for underwater images7
DisRot: boosting the generalization capability of few-shot learning via knowledge distillation and self-supervised learning7
Meta-learning enhanced global–local feature fusion for image quality assessment7
Tensor-guided learning for image denoising using anisotropic PDEs7
Integrating visual-semantic relational reasoning for fake news detection on video platforms7
Automatic cables segmentation from a substation device based on 3D point cloud7
YG-SLAM: dynamic environment-based geometric constraint point-line fusion visual SLAM system7
Improving change detection using conditional discriminative adversarial regularization6
Semi-supervised metric learning incorporating weighted triplet constraint and Riemannian manifold optimization for classification6
Kinematic calibration of a hexapod robot based on monocular vision6
Environmental factors-aware two-stream GCN for skeleton-based behavior recognition6
Cascaded attention-guided multi-granularity feature learning for person re-identification6
Text-to-face synthesis based on facial landmarks prediction6
Accelerated fixed-point iterations for image deblurring and defiltering6
Tree-managed network ensembles for video prediction6
Multi-view dynamic reconstruction with cross-view smoothing based on surfel6
Welding splash and arc noise reduction imaging model based on computationally efficient pairwise response serving welding process library6
Distortion diminishing with vulnerability filters pruning6
Boosting few-shot learning via selective patch embedding by comprehensive sample analysis6
A dual-path U-Net for pulmonary vessel segmentation method based on lightweight 3D attention6
Logit scaling for out-of-distribution detection6
Guest editorial: special issue on human pose estimation and its applications6
A review of adaptable conventional image processing pipelines and deep learning on limited datasets6
PGA6D: 6D pose estimation for grasping and assemblying based on keypoints voting6
Parametric loss-based super-resolution for scene text recognition6
Actions as points: a simple and efficient detector for skeleton-based temporal action detection6
Enhancing adversarial transferability via importance-aware pixel-level mask6
Enhanced normal estimation of point clouds via fine-grained geometric information learning6
Multiple object tracking using weighted graph convolutional neural networks6
Robust semantic segmentation method of urban scenes in snowy environment6
Block-recurrent visual transformer for enhanced human detection in thermal imaging5
PTDS CenterTrack: pedestrian tracking in dense scenes with re-identification and feature enhancement5
Quality assessment of synthetic images via spatial distortion recognition5
Edge-aware dual path network for medical image classification5
Unsupervised single-shot depth estimation using perceptual reconstruction5
Beyond Kalman filters: deep learning-based filters for improved object tracking5
YOLOMH: you only look once for multi-task driving perception with high efficiency5
A parameter-efficient attention-enhanced model for multiscale weld defect detection of wind turbine towers5
Local region-learning modules for point cloud classification5
VGT-MOT: visibility-guided tracking for online multiple-object tracking5
Residual shuffle attention network for image super-resolution5
Regional filtering distillation for object detection5
Human pose estimation based on lightweight basicblock5
Toward phytoplankton parasite detection using autoencoders5
Rid-slam: a robust illumination-and-dynamics-aware RGB-D SLAM framework for indoor environments5
Fine-grained 3D vehicle shape manipulation via latent space editing5
A collaborative SLAM method for dual payload-carrying UAVs in denied environments5
TFF-temporal fusion framework for advancing video retrieval through long-range dependencies and multi-modal intent5
Diffusion-leveraged GAN dehazing driven by classification: a two-stage framework for real-world monitoring imagery5
React: recognize every action everywhere all at once5
BiTransformer: augmenting semantic context in video captioning via bidirectional decoder5
Naturally constrained reject option classification5
Detecting violent deepfakes: dataset and a compact attention network with multi-scale supervision5
Optimized hand pose estimation CrossInfoNet-based architecture for embedded devices4
A lightweight and generalizable detection enhancement method using segmentation feedback4
Vision-based power line cables and pylons detection for low flying aircraft4
Spatial-temporal graph-guided global attention network for video-based person re-identification4
WIFE-Net: widely integrated follicle extraction network4
The general framework for few-shot learning by kernel HyperNetworks4
Symmetry-induced ambiguity in orientation estimation from RGB images4
Normalized margin loss for action unit detection4
ViT-KANMoE: a vision transformer enhanced with a pure Kolmogorov–Arnold network-based mixture-of-experts for malaria blood smear image classification4
Bidirectional cascaded multimodal attention for multiple choice visual question answering4
Token adaptation via side graph convolution for efficient fine-tuning of 3D point cloud transformers4
Enhanced keypoint information and pose-weighted re-ID features for multi-person pose estimation and tracking4
Superpixel-based foreground-preserving image stitching4
Pixel representations, sampling, and label correction for semantic part detection4
Personvit: large-scale self-supervised vision transformer for person re-identification4
ViCap-AD: video caption-based weakly supervised video anomaly detection4
A deep Retinex network for underwater low-light image enhancement4
Self-attention SAC with vision-augmented LiDAR fusion for mapless robot navigation in dynamic environments4
Supervised contrastive learning with multi-scale interaction and integrity learning for salient object detection4
Removing cloud shadows from ground-based solar imagery4
SiamCAR-Kal: anti-occlusion tracking algorithm for infrared ground targets based on SiamCAR and Kalman filter4
A robust vehicle tracking in low-altitude UAV videos4
CCTV-Calib: a toolbox to calibrate surveillance cameras around the globe4
Ipdm: identity preserving diffusion model for face sketch and photo synthesis4
Residual feature learning with hierarchical calibration for gaze estimation4
Ssman: self-supervised masked adaptive network for 3D human pose estimation4
Self-attention network for few-shot learning based on nearest-neighbor algorithm4
Amp: single-shot ultra-wide fisheye-to-cubemap PnP pose estimation4
Swin transformer with part-level tokenization for occluded person re-identification4
Robust 3D object recognition based on improved Hough voting4
A zero-shot anomaly detection method based on learnable text query4
MVUDA: Unsupervised Domain Adaptation for Multi-view Pedestrian Detection4
FLAVR: flow-free architecture for fast video frame interpolation4
Gait recognition using free-area transformer networks4
LS-Occ:light specific-target-focus vision-based 3D occupancy prediction with adaptive combined head4
Multiscale feature optimization for accurate small object detection in remote sensing imagery4
Multimodal dance style transfer4
Carixray: a periapical X-ray dataset for machine vision-based dental caries recognition4
Entangled appearance and motion structures network for multi-object tracking and segmentation4
An image quality assessment method based on edge extraction and singular value for blurriness4
Online camera auto-calibration appliable to road surveillance4
Consensus similarity learning based on tensor nuclear norm3
MYFED: a dataset of affective face videos for investigation of emotional facial dynamics as a soft biometric for person identification3
Virtual home staging and relighting from a single panorama under natural illumination3
Zero-shot action recognition by clustered representation with redundancy-free features3
Interpretability of fingerprint presentation attack detection systems: a look at the “representativeness” of samples against never-seen-before attacks3
Multi-person 3D pose estimation from unlabelled data3
Exploring filter placement in convolutional layer topologies based on ResNet for image classification3
Time-constrained adversarial attacks for video recognition models: temporally sparse but effective perturbations3
Knowledge-based hybrid connectionist models for morphologic reasoning3
Scrap weight prediction for different scrap types based on semantic segmentation and machine learning3
SNFR: salient neighbor decoding and text feature refining for scene text recognition3
Efficient abnormality detection using patch-based 3D convolution with recurrent model3
Similarity contrastive estimation for image and video soft contrastive self-supervised learning3
Wide-baseline multi-camera calibration from a room filled with people3
Generating quality grasp rectangle using Pix2Pix GAN for intelligent robot grasping3
FDT − Dr2T: a unified Dense Radiology Report Generation Transformer framework for X-ray images3
EvoGraphConPain: an adaptive multimodal graph intelligence framework for neonatal pain assessment3
SGBGAN: minority class image generation for class-imbalanced datasets3
CMNet: a novel model and design rationale based on comparison studies and synergy of CNN and MetaFormer3
SFDiff: shadow removal via semantic-guided prototypes and frequency-aware modulation3
Unsupervised domain adaptation by cross-domain consistency learning for CT body composition3
An Efficient point-in-convex 3D polyhedron test using a projective algorithm with sub-linear expected complexity3
Investigating long-term training for remote sensing object detection3
Discriminative feature learning through feature distance loss3
Pixel-wise confidence estimation for segmentation in Bayesian Convolutional Neural Networks3
Exploring the potential of deep learning techniques for analyzing athlete movements in competitive athletics sports3
Material classification of polishing and convex surface objects based on photon accumulation point spread function (PAPSF) from imaging model of binocular pulsed time-of-flight camera3
Interpretable visual transmission lines inspections using pseudo-prototypical part network3
Addressing the generalization of 3D registration methods with a featureless baseline and an unbiased benchmark3
Ising granularity image analysis on VAE–GAN3
FESAR: SAR ship detection model based on local spatial relationship capture and fused convolutional enhancement3
ICE-GCN: An interactional channel excitation-enhanced graph convolutional network for skeleton-based action recognition2
PM-MVS: PatchMatch multi-view stereo2
From explanation to unsupervised segmentation: fusion of multiple explanation maps for vision transformers2
On the effectiveness of MoE-enhanced transformer for accurate and generalizable mask-based semantic segmentation2
Semantic scene upgrades for trajectory prediction2
Uncertainty-guided collaborative learning for noise-robust text-to-image person re-identification2
Editor’s Note: Special Issue on Advances in Visual Computing2
Optimize multiscale feature hybrid-net deep learning approach used for automatic pancreas image segmentation2
Representing dynamic textures based on polarized gradient features2
Biomimetic oculomotor control with spiking neural networks2
Dyna-MSDepth: multi-scale self-supervised monocular depth estimation network for visual SLAM in dynamic scenes2
Multi-scene low-light remote physiological measurement database2
Transformer-based end-to-end multiple object fast-tracking model for golden monkeys2
Tortoise plastron versus adulterants: identification and comparative study using image recognition technology2
Active perception based on deep reinforcement learning for autonomous robotic damage inspection2
Improved deep depth estimation for environments with sparse visual cues2
Pose is all you need: the pose only group activity recognition system (POGARS)2
Editor’s Note: Special Issue on Advances in Visual Computing 20232
Calibrating uncertainties in human trajectory forecasting2
A multi-target physiological signal detection method for UWB radar based on Kalman tracking and dual-branch network2
Self-supervised monocular depth estimation via joint attention and intelligent mask loss2
Multi-core token mixer: a novel approach for underwater image enhancement2
That’s BAD: blind anomaly detection by implicit local feature clustering2
High-efficiency automated triaxial robot grasping system for motor rotors using 3D structured light sensor2
Fast re-OBJ: real-time object re-identification in rigid scenes2
Object Recognition Consistency in Regression for Active Detection2
0.23610401153564