Parallel Computing

Papers
(The TQCC of Parallel Computing is 4. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Editorial Board165
Parallel multi-view HEVC for heterogeneously embedded cluster system29
Heterogeneous sparse matrix–vector multiplication via compressed sparse row format25
Integrating FPGA-based hardware acceleration with relational databases16
A parallel non-convex approximation framework for risk parity portfolio design15
Editorial Board14
NPDP benchmark suite for the evaluation of the effectiveness of automatic optimizing compilers14
LSHDP: Locally sharded heterogeneous data parallel for distributed deep learning13
Mobilizing underutilized storage nodes via job path: A job-aware file striping approach12
Evaluating SYCL as a unified programming model for heterogeneous systems10
Distributed consensus-based estimation of the leading eigenvalue of a non-negative irreducible matrix9
EESF: Energy-efficient scheduling framework for deadline-constrained workflows with computation speed estimation method in cloud8
Task graph-based performance analysis of parallel-in-time methods7
Editorial on Advances in High Performance Programming7
Adaptively parallel runtime verification based on distributed network for temporal properties7
ShyLU-node: On-node scalable solvers and preconditioners: Recent progress and current performance6
ParVoro++: A scalable parallel algorithm for constructing 3D Voronoi tessellations based on kd-tree decomposition6
New YARN sharing GPU based on graphics memory granularity scheduling6
Parallel optimization and application of unstructured sparse triangular solver on new generation of Sunway architecture6
Editorial Board6
OF-WFBP: A near-optimal communication mechanism for tensor fusion in distributed deep learning6
Routing brain traffic through the von Neumann bottleneck: Efficient cache usage in spiking neural network simulation code on general purpose computers6
Multi-level parallelism optimization for two-dimensional convolution vectorization method on multi-core vector accelerator5
Efficient parallel reduction of bandwidth for symmetric matrices5
Editorial Board5
Using Java to create and analyze models of parallel computing systems5
Optimizing convolutional neural networks on multi-core vector accelerator5
A sleek lock-free hash map in an ERA of safe memory reclamation methods4
GPU acceleration of Levenshtein distance computation between long strings4
Editorial Board4
A lightweight semi-centralized strategy for the massive parallelization of branching algorithms4
Targeting performance and user-friendliness: GPU-accelerated finite element computation with automated code generation in FEniCS4
PPS: Fair and efficient black-box scheduling for multi-tenant GPU clusters4
Accelerating the scheduling of the network resources of the next-generation optical data centers4
Tausch: A halo exchange library for large heterogeneous computing systems using MPI, OpenCL, and CUDA4
Byzantine-tolerant detection of causality: There is no holy grail4
A survey of software techniques to emulate heterogeneous memory systems in high-performance computing4
Analyzing the impact of CUDA versions on GPU applications4
0.61639618873596