IEEE Transactions on Circuits and Systems for Video Technology

Papers
(The TQCC of IEEE Transactions on Circuits and Systems for Video Technology is 19. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
2022 Index IEEE Transactions on Circuits and Systems for Video Technology Vol. 32623
IEEE Transactions on Circuits and Systems for Video Technology Publication Information528
Table of Contents372
IEEE Transactions on Circuits and Systems for Video Technology publication information348
IEEE Transactions on Circuits and Systems for Video Technology publication information276
SARGAN: Spatial Attention-Based Residuals for Facial Expression Manipulation257
DMRFlow: 4D Radar Scene Flow Estimation With Decoupled Matching and Refinement253
Table of Contents236
IEEE Circuits and Systems Society Information224
Guest Editorial Introduction to the Special Issue on Label-Efficient Learning on Video Data223
Convolutional Neural Networks for Omnidirectional Image Quality Assessment: A Benchmark197
CRP2-VCS: Contrast-Oriented Region-Based Progressive Probabilistic Visual Cryptography Schemes184
Stochastic Gradient Perturbation: An Implicit Regularizer for Person Re-Identification183
SpiReco: Fast and Efficient Recognition of High-Speed Moving Objects With Spike Camera182
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and User Trajectory Information175
TPCM-SegNet: A Text-Prompted Dual-Path Convolution-Mamba Network for Anomaly Segmentation169
Filtering and Alternating Calibration: Spatiotemporal Context Alternating Fusion for Event-Based Monocular Depth Estimation164
FoV Prediction-Based Adaptive Bitrate Streaming With On-Demand Transcoding for 360° Videos162
LiveMatte: Dynamic Scene Background Restoration and Selective Portrait Patch Enhancement161
Draw Like an Artist: Complex Scene Generation With Diffusion Model via Composition, Painting, and Retouching159
Instance-Incremental Scene Graph Generation From Real-World Point Clouds via Normalizing Flows156
USVTrack: A Benchmark for Multi-Object Tracking in Complex Water Surface Scenes156
Semi-Supervised Crowd Counting via Multi-Task Pseudo-Label Self-Correction Strategy150
UAMD-Net: A Unified Adaptive Multimodal Neural Network for Dense Depth Completion150
Future Feature-Based Supervised Contrastive Learning for Streaming Perception150
Unsupervised Action Segmentation via Multi-Scale Temporal-Interaction Enhancement148
Learning Monocular Depth via Cascaded Iterative Refinement in Visual-Echo Scenes146
Highly-Parallel Hardwired Deep Convolutional Neural Network for 1-ms Dual-Hand Tracking143
MultiHuman: Leverage Multimodal Prompts for Controllable Multi-Person Image Synthesizing142
Relation-Aware Multi-Pass Comparison Deconfounded Network for Change Captioning138
Lightweight Neural Network for Enhancing Imaging Performance of Under-Display Camera135
WordCon: Word-level Typography Control in Visual Text Rendering134
Hierarchical Dynamic Programming Module for Human Pose Refinement133
DS 2 VP: Dynamically-Selected Spatially Visual Prompting132
Enhancing Representation Learning With Spatial Transformation and Early Convolution for Reinforcement Learning-Based Small Object Detection132
Plausible Proxy Mining With Credibility for Unsupervised Person Re-Identification128
EIFNet: An Explicit and Implicit Feature Fusion Network for Finger Vein Verification126
Dual Difficulty-Aware Adaptive Pseudo Labeling for Semi-Supervised CNV Segmentation126
Block Diagonal Graph Embedded Discriminative Regression for Image Representation122
Crowd-Powered Photo Enhancement Featuring an Active Learning Based Local Filter122
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition121
Semantic Boosting via Knowledge Sharing and Feedback for Video Anomaly Detection121
DP-Retinex: Dual-Prior Guided Low-Light Image Enhancement With YUV-Domain Reflectance-Illumination Decomposition121
Push-and-Pull: A General Training Framework With Differential Augmentor for Domain Generalized Point Cloud Classification118
Learning Spatio-Temporal Sharpness Map for Video Deblurring118
MEF-GD: Multimodal Enhancement and Fusion Network for Garment Designer117
Active Spatial Positions Based Hierarchical Relation Inference for Group Activity Recognition116
Key Role Guided Transformer for Group Activity Recognition116
Deep Convolutional Primal-Dual Network for Image Deblurring114
A Clinically Guided Graph Convolutional Network for Assessment of Parkinsonian Pronation-Supination Movements of Hands114
MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression Comprehension111
Subjective and Objective Quality Assessment of Display Content Videos110
Iterative Self-Guided Image Filtering110
FastAL: Fast Evaluation Module for Efficient Dynamic Deep Active Learning Using Broad Learning System109
UDTCWT-PHFMs Domain Statistical Image Watermarking Using Vector BW-Type R Distribution106
NDM: Boosting Dataset Distillation via Nested Difficulty Matching106
ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T Tracking105
Phase-Guided Cross-Frequency Integration Network for ISAR and Optical Image Fusion104
Towards Immersive Volumetric Space: Capture, Reconstruction, and Application103
Multi-Modal Attribute Prompting for Vision-Language Models102
CLIP-Based Class Incremental Semantic Segmentation Framework With Generalization-Preserving Knowledge Distillation102
Scene Prior Constrained Self-Paced Learning for Unsupervised Satellite Video Vehicle Detection101
SMART: Semantic Matching Contrastive Learning for Partially View-Aligned Clustering101
Representing Boundary-Ambiguous Scene Online With Scale-Encoded Cascaded Grids and Radiance Field Deblurring98
VPA: Multi-Modal Virtual Point Augmentation for 3D Object Detection96
Uni3DA: Universal 3D Domain Adaptation for Object Recognition96
Dual-Stream Transformer With Distribution Alignment for Visible-Infrared Person Re-Identification96
TiGDistill-BEV: Multi-View BEV 3D Object Detection via Target Inner-Geometry Learning Distillation95
Towards Video Anomaly Detection in the Real World: A Binarization Embedded Weakly-Supervised Network93
Few-Shot Temporal Sentence Grounding via Memory-Guided Semantic Learning92
Learning Depth-Density Priors for Fourier-Based Unpaired Image Restoration92
Pose-Guided Transformer for Fine-Grained Action Quality Assessment91
Representation Robustness and Feature Expansion for Exemplar-Free Class-Incremental Learning91
Frequency Generation for Real-World Image Super-Resolution91
Exploring and Exploiting High-Order Spatial–Temporal Dynamics for Long-Term Frame Prediction90
Scalable and Robust Tensor Ring Decomposition for Large-Scale Data With Missing Data and Outliers90
Fully Unsupervised Domain-Agnostic Image Retrieval89
Ct-LVI: A Framework Toward Continuous-Time Laser-Visual-Inertial Odometry and Mapping89
Robust Image Watermarking With Synchronization Using Template Enhanced-Extracted Network88
MPCF: Multi-Phase Consolidated Fusion for Multi-Modal 3D Object Detection With Pseudo Point Cloud88
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation88
Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style Mixture88
DSC3D: Deformable Sampling Constraints in Stereo 3D Object Detection for Autonomous Driving88
Morphology-Guided Muscle Cell Detection and Counting Based on Transfer Learning, FFD Augmentation, and Density-Aware Loss Optimization88
Negative Class Guided Spatial Consistency Network for Sparsely Supervised Semantic Segmentation of Remote Sensing Images88
Projected Generative Adversarial Network for Point Cloud Completion87
Boosting Video Object Segmentation With Discriminative Core Features and Adaptive Position Refinement87
Multi-Level Feature Fusion Network for Shadow Removal Detection87
Lossless Dynamic Point Cloud Geometry Compression via Rate-Distortion Optimized Motion Estimation86
RT3DHVC: A Real-Time Human Holographic Video Conferencing System With a Consumer RGB-D Camera Array86
Universal Immunized Cover Construction for Secure Adaptive Steganography Across Multiple Domains86
Multi-Modal Multi-Grained Embedding Learning for Generalized Zero-Shot Video Classification85
Relative Comparison-Based Consensus Learning for Multi-View Subspace Clustering84
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation84
Reconstructing Sparse-View Indoor Scenes in View Space With Global Monocular Prior Alignment83
Local Attention Transformer-Based Full-View Finger-Vein Identification83
Joint Learning of Image Deblurring and Depth Estimation Through Adversarial Multi-Task Network83
Deep and Low-Rank Quaternion Priors for Color Image Processing83
Graph-Guided Unsupervised Multiview Representation Learning83
Equity in Unsupervised Domain Adaptation by Nuclear Norm Maximization83
Reversible Data Hiding Over Encrypted Images via Preprocessing-Free Matrix Secret Sharing82
Spatial Attention-Guided Light Field Salient Object Detection Network With Implicit Neural Representation82
PPIFuse: Physical Priors Injected Infrared and Visible Image Fusion82
HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement81
D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing81
MMI-Det: Exploring Multi-Modal Integration for Visible and Infrared Object Detection81
Spectral–Spatial Feature Extraction With Dual Graph Autoencoder for Hyperspectral Image Clustering80
Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action Recognition80
Exploring Explicitly Disentangled Features for Domain Generalization80
A Format Compliant Framework for HEVC Selective Encryption After Encoding80
Multi-Stage Cross-Modality Feature Interaction for RGB-Thermal Multi-Object Tracking80
AirSOD: A Lightweight Network for RGB-D Salient Object Detection80
ASCFormer: An Adaptive Structure-Aware Cascaded Transformer for 3D Object Detection79
Cross-Level Multi-Modal Features Learning With Transformer for RGB-D Object Recognition77
Efficient Single-Object Tracker Based on Local-Global Feature Fusion77
Learning Appearance-Motion Synergy via Memory-Guided Event Prediction for Video Anomaly Detection76
Dependability Feature Learning Based on Sample Generation for Unsupervised Text-to-Image Person Re-Identification76
Reliable Entropy-Induced Anchor Learning for Incomplete Multi-View Subspace Clustering76
Harmony: An Eco-Friendly Adaptive Rate Control Scheme for Video-on-Demand in Low Earth Orbit Satellite Internet76
VDTR: Video Deblurring With Transformer75
IGDA-ICMH: an Importance Guided Feature and Network Dual Adaptation Framework towards Image Compression for Machine and Human Vision75
Toward Meta-Shape-Based Multi-View 3D Point Cloud Registration: An Evaluation75
Adversarial Dual-Student With Differentiable Spatial Warping for Semi-Supervised Semantic Segmentation75
Video Understanding With Large Language Models: A Survey74
Edge and Skeleton Guidance Network for Salient Object Detection in Optical Remote Sensing Images74
Synergistic Fusion Network of Microscopic Hyperspectral and RGB Images for Multi-Perspective Segmentation74
IEEE Circuits and Systems Society Information74
Pro-Tuning: Unified Prompt Tuning for Vision Tasks74
IEEE Transactions on Circuits and Systems for Video Technology publication information74
Compensating for the Incomplete With the Complete: An Efficient Scene Text Detector73
Multi-Scale Explicit Matching and Mutual Subject Teacher Learning for Generalizable Person Re-Identification73
Inter-Scale Similarity Guided Cost Aggregation for Stereo Matching73
Feature Evaluation and Joint Interaction for Audio-Visual Emotion Recognition73
AMTFusion: boosting 3D object detection by adaptive multi-modal temporal fusion and augmentation73
Mesh2Animation: Unsupervised Animating for Quadruped 3D Objects73
FaceGCN: Structured Priors Inspired Graph Convolutional Networks for Face Restoration With Unknown Degradations73
Blind Image Quality Index for Authentic Distortions With Local and Global Deep Feature Aggregation73
Learning With Noisy Labels by Semantic and Feature Space Collaboration72
Task-Specific Loss for Robust Instance Segmentation With Noisy Class Labels72
Learning Scene-Invariant Distribution for Generalizable Blind Image Quality Assessment72
Texture-Aware Spherical Rotation for High Efficiency Omnidirectional Intra Video Coding72
ImagingNet: A New Learnable SAR Imaging Method via Hierarchical U-Shaped Network72
Table of Contents72
GNNLicense: An Active Intellectual Property Protection Technique for Graph Neural Networks based on Autoencoder71
Folding the Vision: Towards Efficient Global Context Representation on Edge Devices71
Holistic Prototype Attention Network for Few-Shot Video Object Segmentation71
Efficient Non-Blind Image Deblurring With Discriminative Shrinkage Deep Networks70
Table of Contents70
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation70
Target-Aware Tracking With Spatial-Temporal Context Attention70
Errata to “Local-Global Temporal Difference Learning for Satellite Video Super-Resolution”70
Cloth-Imbalanced Gait Recognition via Hallucination70
Non-Local Guided Neural Fields for 4D CT Reconstruction69
Conditional Dual Diffusion for Multimodal Clustering of Optical and SAR Images69
Multimodal Industrial Anomaly Detection via Geometric Prior69
Balanced Teacher for Source-Free Object Detection69
Enhancing Vision Transformer With Shift Expansion Linear Attention for Image Classification and Object Tracking68
Adaptive Mixture-of-Experts Distillation for Cross-Satellite Generalizable Incremental Remote Sensing Scene Classification68
Monocular Depth Estimation on Adverse Weathers With Curriculum Domain Distribution Alignment68
CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning67
FedRSL: Representation Subspace Learning in Model-Heterogeneous Federated Learning67
MDSLCA: Multi-Scale Dilated Spatial and Local Channel Attention for LiDAR Point Cloud Semantic Segmentation67
FDAC: Federated Domain Adaptation via Dual Contrastive Learning67
Self-Supervised Adversarial Video Summarizer With Context Latent Sequence Learning67
One for All: A Unified Generative Framework for Image Emotion Classification67
MSGA-Net: Progressive Feature Matching via Multi-Layer Sparse Graph Attention66
A Label-Free and Non-Monotonic Metric for Evaluating Denoising in Event Cameras66
Diverse Batch Steganography Using Model-Based Selection and Double-Layered Payload Assignment66
An RGB-Event Hierarchical Fusion Enhancement Network for Object Detection66
DEP-Former: Multimodal Depression Recognition Based on Facial Expressions and Audio Features via Emotional Changes66
G2LP-Net: Global to Local Progressive Video Inpainting Network66
All-Inclusive Image Enhancement for Degraded Images Exhibiting Low-Frequency Corruption66
VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality Evaluator66
WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution Diffusion66
MixSSC: Forward-Backward Mixture for Vision-Based 3D Semantic Scene Completion65
Explicit Spatial-Angular Pattern Embedding for Heterogeneous Imaging-Oriented Light Field Spatial Super-Resolution65
Searching a Compact Architecture for Robust Multi-Exposure Image Fusion65
ART-VOD: Adaptive Relational Distillation and Temporal Modeling for Open-Vocabulary Video Object Detection65
DPSFusion: Dynamic Pseudo-Supervision and Semantic Guidance for Multi-Modality Image Fusion64
Question-Aware Global-Local Video Understanding Network for Audio-Visual Question Answering64
Learning Shallow-Robust Vision Transformers with Selective Distillation for Real-time UAV Tracking64
Efficiently Exploiting Spatially Variant Knowledge for Video Deblurring64
OraL: An Observational Learning Paradigm for Unsupervised Hyperspectral Change Detection64
SMR: Spatial-Guided Model-Based Regression for 3D Hand Pose and Mesh Reconstruction63
HyPSAM: Hybrid Prompt-Driven Segment Anything Model for RGB-Thermal Salient Object Detection63
DiffVein: A Unified Diffusion Network for Finger Vein Segmentation and Authentication63
A Universal Framework for Improving the Robustness of Coverless Image Steganography Based on Image Restoration62
Touchless Finger Vein and Fingerprint Verification via Exploiting Attention-Based Cross-Domain Fusion62
Medical Data Security in Blockchain: A Telemedicine Data Sharing Scheme Based on Custom OPE and 4D-YG Hyperchaotic62
CAIR-Net: Reliability-Aware Information Routing for Robust Multimodal Object Detection Under Modality Degradation62
DeGMix: Efficient Multi-Task Dense Prediction with Deformable and Gating Mixer62
Laplacian Pyramid Fusion Network With Hierarchical Guidance for Infrared and Visible Image Fusion61
Flow Visualization for Complex Fluid Flows via a Structure-Enhanced Motion Estimator61
Online Unsupervised Video Object Segmentation via Contrastive Motion Clustering61
YODA: Yet Another One-step Diffusion-based Video Compressor60
Explanation-Guided Adversarial Training for Robust and Interpretable Models60
Unsupervised Deep Hashing With Fine-Grained Similarity-Preserving Contrastive Learning for Image Retrieval60
Optical Flow Reusing for High-Efficiency Space-Time Video Super Resolution59
Enhancing Robustness of Multi-Object Trackers With Temporal Feature Mix59
Continual Semantic Segmentation with Tiny Memory59
Flow-Edge Guided Unsupervised Video Object Segmentation59
STAF: 3D Human Mesh Recovery From Video With Spatio-Temporal Alignment Fusion59
Surveillance Video-and-Language Understanding: From Small to Large Multimodal Models59
Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event Features59
BIMM: Brain-Inspired Masked Modeling for Video Representation Learning59
Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery Detection59
TAKD: Target-Aware Knowledge Distillation for Remote Sensing Scene Classification59
Functionality Separation: Rethinking Dual-Stream Networks for Class-Incremental Learning59
Depth Estimation From a Single Image of Blast Furnace Burden Surface Based on Edge Defocus Tracking58
Enhanced Spatial-Temporal Salience for Cross-View Gait Recognition58
Dynamic Hypergraph Convolutional Network for No-Reference Point Cloud Quality Assessment58
FDNet: Frequency Decomposition Network for Learned Image Compression58
A Novel Deep Learning Framework for Automatic Recognition of Thyroid Gland and Tissues of Neck in Ultrasound Image58
Integrating Pseudo-Supervision and Spatial Constraints for Efficient Clustering of Multimodal Remote Sensing Data58
Meta-Learning Based Domain Prior With Application to Optical-ISAR Image Translation58
Semantic Disentanglement Adversarial Hashing for Cross-Modal Retrieval57
Low-Resolution Object Recognition With Cross-Resolution Relational Contrastive Distillation57
Robust Matrix Completion Based on Factorization and Truncated-Quadratic Loss Function57
CNN-Transformer Based Generative Adversarial Network for Copy-Move Source/ Target Distinguishment57
WeakLoc: Low-Rank Attention and Adaptive Pseudo-Label Learning for Weakly Supervised Image Forgery Localization57
A Physical Model-Guided Framework for Underwater Image Enhancement and Depth Estimation57
Recent Advances in Rate Control: From Optimization to Implementation and Beyond57
StreetSurfGS: Scalable Urban Street Surface Reconstruction With Planar-Based Gaussian Splatting57
Multi-Prior Driven Network for RGB-D Salient Object Detection57
Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the Wild57
M3CS: Multi-Target Masked Point Modeling With Learnable Codebook and Siamese Decoders56
Table of Contents56
IEEE Transactions on Circuits and Systems for Video Technology publication information56
VmambaIR: Visual State Space Model for Image Restoration56
Prototype Decoupled Knowledge Distillation55
Concept-Enhanced Relation Network for Video Visual Relation Inference55
Dual-Path Feature Aware Network for Remote Sensing Image Semantic Segmentation55
Knowledge-Based Visual Question Generation54
Reference-Guided Large-Scale Face Inpainting With Identity and Texture Control54
Revisiting Modality-Specific Feature Compensation for Visible-Infrared Person Re-Identification54
Bi-Directional Progressive Guidance Network for RGB-D Salient Object Detection54
U²-Former: Nested U-Shaped Transformer for Image Restoration via Multi-View Contrastive Learning54
Flexible Temperature Parallel Distillation for Dense Object Detection: Make Response-Based Knowledge Distillation Great Again54
Real Image Denoising via Guided Residual Estimation and Noise Correction54
Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object Detection54
Neuron-Based Spiking Transmission and Reasoning Network for Robust Image-Text Retrieval54
Improving Zero-Shot Generalization for CLIP With Prompt Ensemble Self-Distillation53
Exploring Implicit Domain-Invariant Features for Domain Adaptive Object Detection53
Deep Video Super-Resolution Using Hybrid Imaging System53
Generalized Intra-Camera Supervised Person Re-Identification53
Corruption-Invariant Person Re-Identification via Coarse-to-Fine Feature Alignment53
Special Issue on Segment Anything for Videos and Beyond53
SPCL: Semantic Polymorphism and Commonality Learning for Text-Based Person Retrieval53
Zero-shot Segmentation with Co-occurrence Relation Enhancement53
PVF-DectNet++: Adaptive Multi-Modal Fusion With Perspective Voxels for 3D Object Detection53
0.48428583145142