Parallel Computing

Papers
(The median citation count of Parallel Computing is 1. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Editorial Board165
Parallel multi-view HEVC for heterogeneously embedded cluster system29
Heterogeneous sparse matrix–vector multiplication via compressed sparse row format25
Integrating FPGA-based hardware acceleration with relational databases16
A parallel non-convex approximation framework for risk parity portfolio design15
Editorial Board14
NPDP benchmark suite for the evaluation of the effectiveness of automatic optimizing compilers14
LSHDP: Locally sharded heterogeneous data parallel for distributed deep learning13
Mobilizing underutilized storage nodes via job path: A job-aware file striping approach12
Evaluating SYCL as a unified programming model for heterogeneous systems10
Distributed consensus-based estimation of the leading eigenvalue of a non-negative irreducible matrix9
EESF: Energy-efficient scheduling framework for deadline-constrained workflows with computation speed estimation method in cloud8
Task graph-based performance analysis of parallel-in-time methods7
Editorial on Advances in High Performance Programming7
Adaptively parallel runtime verification based on distributed network for temporal properties7
ShyLU-node: On-node scalable solvers and preconditioners: Recent progress and current performance6
ParVoro++: A scalable parallel algorithm for constructing 3D Voronoi tessellations based on kd-tree decomposition6
New YARN sharing GPU based on graphics memory granularity scheduling6
Parallel optimization and application of unstructured sparse triangular solver on new generation of Sunway architecture6
Editorial Board6
OF-WFBP: A near-optimal communication mechanism for tensor fusion in distributed deep learning6
Routing brain traffic through the von Neumann bottleneck: Efficient cache usage in spiking neural network simulation code on general purpose computers6
Multi-level parallelism optimization for two-dimensional convolution vectorization method on multi-core vector accelerator5
Efficient parallel reduction of bandwidth for symmetric matrices5
Editorial Board5
Using Java to create and analyze models of parallel computing systems5
Optimizing convolutional neural networks on multi-core vector accelerator5
A sleek lock-free hash map in an ERA of safe memory reclamation methods4
GPU acceleration of Levenshtein distance computation between long strings4
Editorial Board4
A lightweight semi-centralized strategy for the massive parallelization of branching algorithms4
Targeting performance and user-friendliness: GPU-accelerated finite element computation with automated code generation in FEniCS4
PPS: Fair and efficient black-box scheduling for multi-tenant GPU clusters4
Accelerating the scheduling of the network resources of the next-generation optical data centers4
Tausch: A halo exchange library for large heterogeneous computing systems using MPI, OpenCL, and CUDA4
Byzantine-tolerant detection of causality: There is no holy grail4
A survey of software techniques to emulate heterogeneous memory systems in high-performance computing4
Analyzing the impact of CUDA versions on GPU applications4
Lifeline-based load balancing schemes for Asynchronous Many-Task runtimes in clusters3
Enable cross-iteration parallelism for PIM-based graph processing with vertex-level synchronization3
Random sketching to enhance the numerical stability of block orthogonalization algorithms for s-step GMRES3
Distributed software defined network-based fog to fog collaboration scheme3
Butterfly factorization for vision transformers on multi-IPU systems3
Editorial Board3
A flexible sparse matrix data format and parallel algorithms for the assembly of finite element matrices on shared memory systems3
Exploring metrics for analyzing dynamic behavior in MPI programs via a coupled-oscillator model3
Spatial-aware data partition for distributed memory parallelization of ANN search in multimedia retrieval3
Editorial Board3
Parallel Pattern Compiler for Automatic Global Optimizations2
QoS-aware dynamic resource allocation with improved utilization and energy efficiency on GPU2
Cache partitioning for sparse matrix–vector multiplication on the A64FX2
LSAF: A load-balancing SpGEMM acceleration framework with dynamic package and static partition for multi-core systolic arrays2
FastPTM: Fast weights loading of pre-trained models for parallel inference service provisioning2
Editorial Board2
NekRS, a GPU-accelerated spectral element Navier–Stokes solver2
Editorial for parallel computing2
SGPM: A coroutine framework for transaction processing1
Accelerating communication for parallel programming models on GPU systems1
Optimizing massively parallel sparse matrix computing on ARM many-core processor1
WBSP: Addressing stragglers in distributed machine learning with worker-busy synchronous parallel1
GPU/CUDA-Accelerated gradient growth optimizer for efficient complex numerical global optimization1
Performance and accuracy predictions of approximation methods for shortest-path algorithms on GPUs1
Editorial Board1
HRPF: A parallel programming framework for recursive algorithms on heterogeneous CPU–GPU systems1
An optimal scheduling algorithm considering the transactions worst-case delay for multi-channel hyperledger fabric network1
Operational Data Analytics in practice: Experiences from design to deployment in production HPC environments1
ALBBA: An efficient ALgebraic Bypass BFS Algorithm on long vector architectures1
Multi-level parallel multi-layer block reproducible summation algorithm1
Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers1
Fast calculation of isostatic compensation correction using the GPU-parallel prism method1
A heterogeneous processing-in-memory approach to accelerate quantum chemistry simulation1
A survey of parallel computing frameworks and optimizations for AI and deep learning1
Benchmark of classical disk array and software-defined storage on near-identical hardware1
Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid1
Big data BPMN workflow resource optimization in the cloud1
An approach for low-power heterogeneous parallel implementation of ALC-PSO algorithm using OmpSs and CUDA1
Extending the limit of LR-TDDFT on two different approaches: Numerical algorithms and new Sunway heterogeneous supercomputer1
Analysis of the impact of NUMA node configuration on the performance of offloading computations to GPUs1
0.046442031860352