Multimedia Systems

Papers
(The TQCC of Multimedia Systems is 4. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Pseudo-global strategy-based visual comfort assessment considering attention mechanism133
SS-CMT: a label independent cross-modal transferable adversarial video attack with sparse strategy98
Face and voice cross-modal association with learning convex feature embedding70
DiffRA: universal restorative adversarial attack based on diffusion model70
TreeSegNet: multi-scale query-based instance segmentation with frequency-aware and gated feature enhancement57
Dual-branch spectral–spatial feature extraction network for multispectral image compression56
A research for sound event localization and detection based on local–global adaptive fusion and temporal importance network53
FedMAB: adaptive multimodal federated learning with multi-armed bandits50
A visual question answering model based on image captioning47
Unsupervised deep metric learning algorithm for crop disease images based on knowledge distillation networks45
On-line monitoring of structural performance of scraper conveyor driven by digital twin44
Multi-view Isolated sign language recognition based on cross-view and multi-level transformer44
Model-based portrait video compression with spatial constraint and adaptive pose processing43
JAMD-Net: image splicing forgery detection based on JPEG compression artifacts and multi-dilated channel refinement fusion39
Segmentation-aware image super-resolution with generative adversarial networks36
CHCoT-MSLU: a coupled hierarchical chain-of-thought prompt learning model for multi-intent spoken language understanding34
Real emotion seeker: recalibrating annotation for facial expression recognition33
Towards domain adaptation underwater image enhancement and restoration32
Fast latent-feature augmentation for cross-domain face forgery detection32
SFRA: spatial fusion regression augmentation network for facial landmark detection32
A comparative study of color quantization methods using various image quality assessment indices31
Feature fusion and optimization integrated refined deep residual network for diabetic retinopathy severity classification using fundus image31
360° video quality assessment based on saliency-guided viewport extraction31
LEA-depth: a lightweight self-supervised monocular depth estimation with attention fusion and edge-aware distillation30
GVA: guided visual attention approach for automatic image caption generation30
The segmented UEC Food-100 dataset with benchmark experiment on food detection30
Mamba-driven context-aware tracking with dual prompts27
GCGV: a dual-branch hybrid network integrating graph attention, CNNs, and vision transformers for enhanced hyperspectral image classification27
BENet: bi-directional enhanced network for image captioning26
ConASD: Contrastive Few Shot Learning for Detecting Autism Spectrum Disorder via Eye Tracking Scanpath26
Dual convolutional neural network with attention for image blind denoising26
SEMNet: a simple and efficient MLP-based network for 3D Face point clouds landmarks localization26
Hierarchical feature multi-contrastive learning for skin cancer classification25
Atacr-net: adaptive temporal alignment and contrastive refinement network for skeleton-based action recognition25
Saliency guided deep unfolding network for compressive sensing25
Semi-supervised adversarial training via disentangled contrastive learning25
Sketch-guided neural style transfer24
Automatic lymph node segmentation using deep parallel squeeze & excitation and attention Unet24
Rating-aware argument generation for movie reviews with multimodal large language models and a new dataset23
Enhancing person re-identification with gait silhouettes23
A variational causal inference-based method for recognizing object state changes in videos23
Generalizing sentence-level lipreading to unseen speakers: a two-stream end-to-end approach23
Design and realization of pulse-controlled multi-memristor Hopfield neural networks and their applications in information encryption22
CAPNet: tomato leaf disease detection network based on adaptive feature fusion and convolutional enhancement22
Multi-level sentiment-aware clustering for denoising in multimodal sentiment analysis with ASR errors21
User authentication method based on keystroke dynamics and mouse dynamics using HDA21
Enhancing feature diversity with weak teacher directed parameters perturbation21
SS-YOLOv8: small-size object detection algorithm based on improved YOLOv8 for UAV imagery20
Optimizing codebook training through control chart analysis20
Game and reference: efficient policy making for epidemic prevention and control20
Deep Learning-based forgery detection and localization for compressed images using a hybrid optimization model20
SFFN-YOLO for small object detection in aerial images20
Fast bilateral filter with spatial subsampling20
LMFE-RDD: a road damage detector with a lightweight multi-feature extraction network20
SoftBinReduce: data reduction for color quantization through soft binning20
Spatial interpolation of head-related transfer functions using a physics-informed autoencoder20
Big-LITTLE-Net: a dual-branch network for small UAV detection19
CGMAformer: CNN and gated multi axial-sparse transformer feature fusion network for image deraining19
A verifiable variable threshold visual image secret sharing scheme19
Inter-class distance enhanced prototypical network for few-shot text classification19
RefinerHash: a new hashing-based re-ranking technique for image retrieval18
An automatic music generation method based on RSCLN_Transformer network18
Quantifying Factual Divergence in Generative Models: SHAP-LIME Based Hallucination Score for LLMs18
Enhanced target recognition and localization using binocular vision and infrared thermal imaging18
Incrementaldreamer: scene-level 3D generation with incremental optimization18
DMFTNet: dense multimodal fusion transfer network for free-space detection17
Similarity-guided contrastive learning for deep multi-view clustering17
Pull and concentrate: improving unsupervised semantic segmentation adaptation with cross- and intra-domain consistencies17
A multi-label classification method combined with texture enhancement for deepfake face detection17
Speech-driven talking face video generation17
Vulnerability Positioner (VulP): enhancing code vulnerability localization with CodeBERT16
DAFMixerSR: a lightweight fusion-enhanced adaptive perception network for image super-resolution16
EDB-Diff: a EdgeDevice based diffusion network for brain tumor image segmentation16
Exploiting local detail in single image super-resolution via hypergraph convolution16
Design and evaluation of a serious game in virtual reality to increase empathy towards students with phonological dyslexia16
RGB-Net: transformer-based lightweight low-light image enhancement network via RGB channel separation16
Transferable diffusion transformer for low-light image enhancement16
Fgef-net: frequency-guided and enhanced fusion dehazing network for visibility enhancement in maritime traffic surveillance16
Bcgn: BLIP-based cross-modal grasping network for language-conditioned robotic grasping16
GL-MambaNet: Mamba-based global and local feature fusion for image dehazing16
CCM-Net: image splicing localization network based on context-aware and cross-domain multi-scale fusion16
Efficient image-text retrieval via bi-cross-graph learning and multi-grained alignment15
Badinterpreter: Backdoor attack on LLM-based interpretable recommendation15
Learning complementary features via cross-modal attention for robust 6-DoF grasping15
VLM-driven fine-grained semantic regularization for low-light image enhancement15
Weakly supervised anomaly detection with multi-level contextual modeling14
A survey of multimodal federated learning: background, applications, and perspectives14
Diffusion-based synthetic rating generation to alleviate data sparsity in recommender systems14
CR-DM: A novel craniofacial reconstruction framework based on diffusion model14
Multi-view region proposal network predictive learning for tracking14
Enhanced 3D reconstruction with all-neighbor-first philosophy and Ricci flow-based mesh smoothing approach13
Graph-CFRNet: contextual fusion refinement for multi-agent trajectory prediction in autonomous driving13
CMLCNet: medical image segmentation network based on convolution capsule encoder and multi-scale local co-occurrence13
Skeleton-based human activity recognition with wifi CSI using a hybrid approach combining convolutional neural network and long short term memory13
PDSRN: a progressive distillation network for generalizable single image super-resolution13
EDCM-EA: event prediction based on event development context mining considering event arguments13
AI-driven Braille character recognition using partitioned spatial modeling and sequential learning13
Semantic segmentation network for remote sensing images based on category-aware cross-fusion13
Occluded scene text detection via context-awareness from sketch-level image representations13
Workpiece tracking based on improved SiamFC++ and virtual dataset13
Cross-modality geometry-guided historical momentum learning for coupled noisy visible-infrared re-identification13
Object detection of mural images based on improved YOLOv813
NDAM-YOLOseg: a real-time instance segmentation model based on multi-head attention mechanism13
Style matching CAPTCHA: match neural transferred styles to thwart intelligent attacks13
Computer-aided diagnosis for early detection and staging of human pancreatic tumors using an optimized 3D CNN on computed tomography12
Learning shared features from specific and ambiguous descriptions for text-based person search12
Rethinking RGB-D salient object detection12
Enhancing long-tailed classification via multi-strategy weighted experts with hybrid distillation12
UAPT: an underwater acoustic target recognition method based on pre-trained Transformer12
Depth alignment interaction network for camouflaged object detection12
3D human pose estimation method based on multi-constrained dilated convolutions12
A plug-and-play image enhancement model for end-to-end object detection in low-light condition12
Graph contrastive learning for recommendation with generative data augmentation12
Scd-yolo: a novel object detection method for efficient road crack detection12
PCAF: UAV scenarios detector via pyramid converge-and-assign fusion network12
Multimodal-enhanced hierarchical attention network for video captioning12
Joint $$\alpha {-}\beta $$-divergences reconstruction and non-convex sparse regularization for image clustering12
A CNN-transformer hybrid network with selective fusion and dual attention for image super-resolution12
MCLSC-Fusion: a multi-scale cross-modality long-short connection fusion network for infrared and visible images12
Federated learning for medical image classification based on prototype alignment12
FFAA-Net: a full-scale frequency-aware and anisotropic attention network for UAV object detection11
LPR: learning point-level temporal action localization through re-training11
A comprehensive survey on human pose estimation approaches11
HSGNet: hierarchically stacked graph network with attention mechanism for 3D human pose estimation11
Unsupervised knowledge representation of panoramic dental X-ray images using SVG image-and-object clustering11
AF-MT: adaptive fusion based on mean teacher with alternating loss update for medical image segmentation11
An efficient video quality assessment model incorporating VideoMamba and dual-dimensional attention11
LAM-YOLOv11 for UAV transmission line inspection: overcoming environmental challenges with enhanced detection efficiency11
Gicnet: global information capture network for visual place recognition11
Facial action unit detection with emotion consistency: a cross-modal learning approach11
Tb-mmrd: transformer-based multi-modal election rumor detection with agreement-aware gating and semantic fusion11
Segmentation-guided transformer network for subtle visual relationship detection11
Multi-domain feature enhanced adaptive fusion network for multi-modal fake news detection11
Overcomplete-to-sparse representation learning for few-shot class-incremental learning11
Multi-level fine-grained center calibration network for unsupervised person re-identification11
DFGAnet: a dual-branch multimodal fusion network based on graph and attention for emotion recognition in conversation11
Swiftavatar: real-time human reconstruction via semantic graph deformation and surface awareness10
Pointlgfn: local–global fusion network for point cloud classification10
SADCL-Net: Sparse-driven Attention with Dual-Consistency Learning Network for Incomplete Multi-view Clustering10
Remote sensing image cloud removal based on multi-scale spatial information perception10
Deepfake detection of occluded images using a patch-based approach10
3D model watermarking using surface integrals of generated random vector fields10
Diff-mednet: differential convolution and median-enhanced attention multiscale fusion for infrared small target detection10
Polarity-aware attention network for image sentiment analysis10
Correction: Multi-scale prototype contrast and feature fusion for visible-infrared person re-identification10
PMFS-FLC: a partial multi-label feature selection with feature and label collaboration10
COVID-SegNet: encoder–decoder-based architecture for COVID-19 lesion segmentation in chest X-ray10
You watch once more: a more effective CNN architecture for video spatio-temporal action localization10
Gbp-llm: gaze behavior prediction in 6DoF VR via large language models10
Attention mechanisms in deep learning for surface lesion diagnosis: a comprehensive review10
CloudCap3D: enhancing 3D in-scene descriptions via point cloud integration and efficient text filtering9
Learning unified anchor graph based on affinity relationships with strong consensus for multi-view spectral clustering9
Hfffap-net: unsupervised fundus image enhancement with high-frequency feature fusion and artifact processing9
A Three-stage multimodal emotion recognition network based on text low-rank fusion9
Same-clothes person re-identification with dual-stream network9
Unsupervised single-image dehazing via self-guided inverse-retinex GAN9
Dual-stream progressive neural network based on cross fusion in image manipulation localization9
Dual-visual collaborative enhanced transformer for image captioning9
A hybrid spatial and spectral mamba network for hyperspectral image super-resolution9
Panoramic image semantic segmentation using channel attention-based HarDNet and distorted boundary learning9
GloFP-MSF: monocular scene flow estimation with global feature perception9
Tex-Net: texture-based parallel branch cross-attention generalized robust Deepfake detector9
Face attribute recognition via end-to-end weakly supervised regional location9
Task-adaptive parameter optimization for medical image classification transfer learning9
ReDiT: re-evaluating large visual question answering model confidence by defining input scenario difficulty and applying temperature mapping9
MobileViNeXt: a lightweight fusion model for ship-radiated noise recognition9
CBLC-SOOD: contrastive background and label correction for semi-supervised oriented object detection9
Text-centered cross-sample fusion network for multimodal sentiment analysis9
Non-convex fractional-order TV model for image inpainting9
Synthetic shadows: the interplay of forensic detection and anti-forensic techniques in GAN-generated images9
FedVC-ADDiM: a federated learning framework for diagnosis of alzheimer disease using deep learning9
Reducing blind spots in esophagogastroduodenoscopy examinations using a novel deep learning model9
3D human pose estimation with multi-hypotheses gated transformer9
ST-GRU: spatiotemporal gated recurrent unit for video prediction9
Fine-grained behavior interaction-aware network for efficient multi-person motion forecasting9
Multi-object tracking in the low-light with two-stage association and denoising based on image feature enhancement9
A robust federated aggregation algorithm for multimodal data in smart grid scenarios9
Compact twice fusion network for edge detection9
Bag of states: a non-sequential approach to video-based engagement measurement9
Detecting offensive language on instagram with a combined approach of the Gray Wolf algorithm and deep learning networks9
Dual attention transformer with adaptive frequency enhancement for real-world Chinese–English scene text image super-resolution8
EA-EDNet: encapsulated attention encoder-decoder network for 3D reconstruction in low-light-level environment8
Accurate pixel-wise keypoint localization for rectangle symbol spotting in CAD images8
Hierarchical MVSNet with cost volume separation and fusion based on U-shape feature extraction8
Multi-document localization method based on bottom-up architecture8
Adversarial training in logit space against tiny perturbations8
Scene text image super-resolution algorithm based on directional feature modeling8
Attention-guided multi-scale weakly supervised learning for tiny pest detection in citrus orchards8
VCounselor: a psychological intervention chat agent based on a knowledge-enhanced large language model8
DSFusion: different size modalities zero-shot segmentation via heterogeneous fusion8
Advanced techniques in digital media processing for special effects enhancement in film and television post-production8
Multi-granular dynamic interaction network for multimodal sarcasm detection8
Teaching authentic sign language through multiple representation learning8
Hybrid embedding for multimodal few-frame action recognition8
KECAN: knowledge-enhanced cross-modal alignment network for ophthalmic report generation8
DyG-HSTA: a dynamic graph hybrid spatio-temporal attention model for Ethereum phishing detection8
Gmd: Gaussian mixture descriptor for pair matching of 3D fragments8
Local discriminative graph convolutional networks for text classification8
Opfusion: a deep blind image super resolution network using generative diffusion models and neural operator learning8
Cigfuse: an infrared and visible image fusion algorithm based on cross-modal interaction and gating mechanism8
MWEFDet: mamba and wavelet enhanced fusion toward multispectral object detection8
EfficientFace: an efficient deep network with feature enhancement for accurate face detection8
Diving performance analysis with 3D motion knowledge hypergraphs8
DS-Diff: a dual-stage network with degradation-aware and semantic-aware for adverse weather removal based on diffusion models8
ASFESRN: bridging the gap in real-time corn leaf disease detection with image super-resolution8
Selecting generated synthetic features using clustering algorithm for generalized zero-shot learning8
Prior-based bi-encoder transformer for underwater image enhancement8
Dual-guided multi-modal bias removal strategy for temporal sentence grounding in video8
Hierarchical segmentation for traditional cultural pattern based on iterative compression and clustering8
Student engagement detection in online environment using computer vision and multi-dimensional feature fusion8
Generating generalized zero-shot learning based on dual-path feature enhancement8
Point cloud upsampling with implicit graph neural networks8
Lightweight super-resolution via multi-group window self-attention and residual blueprint separable convolution8
GCMR-Net: A Global Context-Enhanced Multi-scale Residual Network for medical image segmentation8
Gender estimation based on deep learned and handcrafted features in an uncontrolled environment8
MGSAN: multimodal graph self-attention network for skeleton-based action recognition8
STSD: spatial–temporal semantic decomposition transformer for skeleton-based action recognition8
PAR-mono: monocular video depth estimation network based on channel separation and dynamic attention8
Composite makeup transfer model based on generative adversarial networks7
Hierarchical segmentation-guided diffusion framework for high-fidelity sonar image generation7
DRL-based transmission control for QoE guaranteed transmission efficiency optimization in tile-based panoramic video streaming7
Estimating visibility via differential regression network7
Lightweight dual-path octave generative adversarial networks for few-shot image generation7
Df-emii: a dual-level multimodal fusion framework for sentiment analysis and emotion recognition with applications in public opinion monitoring7
Multiscale geometric window transformer for orthodontic teeth point cloud registration7
Msfusenet: a multi-stage information fusion network for multi-modal skin lesion diagnosis7
A prompt-based dual-layer cross-modal distillation learning method for aspect-based sentiment analysis7
Gated feature aggregate and alignment network for real-time semantic segmentation of street scenes7
TIPDF-DWSF: a task-oriented two-stage optimization framework for diffusion model LoRA fine-tuning7
Kronecker-factored Approximate Curvature with adaptive learning rate for optimizing model-agnostic meta-learning7
Propagating prior information with transformer for robust visual object tracking7
Fast-colorfool: faster and more transferable semantic adversarial attack with complementary colors and cumulative perturbation7
Blind super-resolution based on matrix-variable optimization for video images7
Recognition of miner action and violation behavior based on the ANODE-GCN model7
Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS7
Wavelet guided real time detection transformer with sparse attention7
SR-DAYOLOv8: cross-domain adaptive object detection based on super-resolution domain classifier7
Rescue decision via Earthquake Disaster Knowledge Graph reasoning7
TSGFormer: temporal-aware network and spatial encoding GCN for three-dimensional human pose estimation7
CAFIN: cross-attention based face image repair network7
Role of deep learning models and analytics in industrial multimedia environment7
Link prediction in social networks using hyper-motif representation on hypergraph7
Fine-tuning CLIP for difference-guided composed image retrieval7
DADP: a discrepancy-aware dual-path network for AI-generated image detection7
Deep unfolding low-rank network for image denoising7
YOLO-ERF: lightweight object detector for UAV aerial images7
Channel modulus normalization for CNN image classification7
Personalized time-sync comment generation based on a multimodal transformer7
An adaptive Bagging algorithm based on lightweight transformer for multi-class imbalance recognition7
Context-aware feature complementary screening network for mass segmentation in whole mammograms7
0.55813503265381