IEEE Transactions on Parallel and Distributed Systems

Papers
(The H4-Index of IEEE Transactions on Parallel and Distributed Systems is 51. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Critique of “MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization” by SCC Team From Tsinghua University209
Design and Implementation of 2D Convolution on x86/x64 Processors208
Distributed Task Processing Platform for Infrastructure-Less IoT Networks: A Multi-Dimensional Optimization Approach163
H5Intent: Autotuning HDF5 With User Intent150
AWB+-Tree: A Novel Width-Based Index Structure Supporting Hybrid Matching for Large-Scale Content-Based Pub/Sub Systems143
fPIM: A Holistic Design to Optimize PIM Data Flow for High Execution Efficiency124
Fully Decentralized Data Distribution for Large-Scale HPC Systems119
Mapping Large-Scale Spiking Neural Network on Arbitrary Meshed Neuromorphic Hardware118
STR: Hybrid Tensor Re-Generation to Break Memory Wall for DNN Training114
Replicated Versioned Data Structures for Wide-Area Distributed Systems110
On the Message Complexity of Fault-Tolerant Computation: Leader Election and Agreement109
A Memory-Constraint-Aware List Scheduling Algorithm for Memory-Constraint Heterogeneous Muti-Processor System103
mtGEMM: An Efficient GEMM Library for Modern Multi-Core DSPs100
Jdebug: A Fast, Non-Intrusive and Scalable Fault Locating Tool for Ten-Million-Scale Parallel Applications99
Large-Scale Neural Network Quantum States Calculation for Quantum Chemistry on a New Sunway Supercomputer90
A Point Cloud Video Recognition Acceleration Framework Based on Tempo-Spatial Information89
UniOrch: A Unified Mixed Framework for High-Efficiency LLM Training on Heterogeneous AI Chips86
Bal-DGCN: A Hardware Acceleration Framework for Balanced Computational Efficiency in DGCNs85
Optimizing Data Locality by Integrating Intermediate Data Partitioning and Reduce Task Scheduling in Spark Framework84
Decentralized Federated Learning With Period Gradient Tracking Over Time-Varying Networks79
Federated Learning With Nesterov Accelerated Gradient78
HRCM: A Hierarchical Regularizing Mechanism for Sparse and Imbalanced Communication in Whole Human Brain Simulations78
Enabling Large Scale Simulations for Particle Accelerators76
IRHunter: Universal Detection of Instruction Reordering Vulnerabilities for Enhanced Concurrency in Distributed and Parallel Systems74
ComStar: Compression-Aware Stream Query for Heterogeneous Hybrid Architecture72
GeoScale: Microservice Autoscaling With Cost Budget in Geo-Distributed Edge Clouds71
QoS-Aware Scheduling of Remote Rendering for Interactive Multimedia Applications in Edge Computing70
Online Container Caching for IoT Data Processing in Serverless Edge Computing69
An Efficient Bottleneck Planes Exclusion Method for Reconfiguring 3D VLSI Arrays67
Sparsity-Aware Fault Tolerance for Reliable CNN Inference on GPUs67
RHINO: An Efficient Serverless Container System for Small-Scale HPC Applications66
EdgeTB: A Hybrid Testbed for Distributed Machine Learning at the Edge With High Fidelity66
DyLaClass: Dynamic Labeling Based Classification for Optimal Sparse Matrix Format Selection in Accelerating SpMV65
HarmonyCache: Scalable In-Network Cache With Read-Write Separation65
PHIDE: A Parallel Hybrid Direct–Iterative Eigensolver for Hermitian Eigenvalue Problems63
On the Performance of SMASH: A Non-Preemptive Window-Based Scheduler for Multiserver Jobs63
Tag-Sharer-Fusion Directory: A Scalable Coherence Directory With Flexible Entry Formats62
BARM: A Batch-Aware Resource Manager for Boosting Multiple Neural Networks Inference on GPUs With Memory Oversubscription62
Graph-Centric Performance Analysis for Large-Scale Parallel Applications62
A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration60
Scalable Hybrid Learning Techniques for Scientific Data Compression59
Efficient and Automated Deployment Architecture for OpenStack in TianHe SuperComputing Environment58
Building Accurate and Interpretable Online Classifiers on Edge Devices56
Cannikin: No Lagger of SLO in Concurrent Multiple LoRA LLM Serving55
PreTrans: Enabling Efficient CGRA Multi-Task Context Switch Through Config Pre-Mapping and Data Transceiving54
AESM2 Attribute-Based Encrypted Search for Multi-Owner and Multi-User Distributed Systems54
Multi-Swarm Co-Evolution Based Hybrid Intelligent Optimization for Bi-Objective Multi-Workflow Scheduling in the Cloud53
Simple, Fast and Widely Applicable Concurrent Memory Reclamation via Neutralization52
CiMBA: Accelerating Genome Sequencing Through On-Device Basecalling via Compute-in-Memory52
Adaptive Image Batching and Slicing in Edge Networks for Delay-Critical Small Object Detection51
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network51
0.1030969619751