ACM Transactions on Architecture and Code Optimization

Papers
(The median citation count of ACM Transactions on Architecture and Code Optimization is 1. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
TNT: A Modular Approach to Traversing Physically Heterogeneous NOCs at Bare-wire Latency62
ASM: An Adaptive Secure Multicore for Co-located Mutually Distrusting Processes52
ESMPC: An Efficient Neural Network Training Framework for Secure Two- and Three-Party Computation38
Accelerating Verifiable Queries over Blockchain Database System Using Processing-in-memory33
Intra-request Lag-aware Cache Management to Enhance I/O Responsiveness of SSDs31
Supporting QoS Guarantee in Heterogeneous Object Storage System: A Spatio-Temporal Graph Data Processing Method30
Highly Efficient Self-checking Matrix Multiplication on Tiled AMX Accelerators30
An Intelligent Scheduling Approach on Mobile OS for Optimizing UI Smoothness and Power29
Performance, Energy and NVM Lifetime-Aware Data Structure Refinement and Placement for Heterogeneous Memory Systems29
HierMine: Accelerating Graph Pattern Mining via Hierarchical Sampling27
Characterizing Digital DRAM PIM through Modeling and Benchmarking26
TransCL: An Automatic CUDA-to-OpenCL Programs Transformation Framework25
ModNEF : An Open Source Modular Neuromorphic Emulator for FPGA for Low-Power In-Edge Artificial Intelligence25
Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary Translation22
A Concise Concurrent B + -Tree for Persistent Memory20
Source Matching and Rewriting for MLIR Using String-Based Automata18
DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor18
COER: A Network Interface Offloading Architecture for RDMA and Congestion Control Protocol Codesign18
Fast Convolution Meets Low Precision: Exploring Efficient Quantized Winograd Convolution on Modern CPUs17
Enabling Efficient All-Dimension Top-k Sparsification for High-Performance Distributed DNN Training Systems15
DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster14
Osiris: A Systolic Approach to Accelerating Fully Homomorphic Encryption14
iSwap: A New Memory Page Swap Mechanism for Reducing Ineffective I/O Operations in Cloud Environments14
Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise Product13
A NUMA-Aware Version of an Adaptive Self-Scheduling Loop Scheduler13
FlashGEMM: Optimizing Sequences of Matrix Multiplication by Exploiting Data Reuse on CPUs12
PredComp: Predicting Compiler Optimization Options with Multi-stage Learning12
SnsBooster: Enhancing Sampling-based μ Arch Evaluation Efficiency through Online Performance Sensitivity Analysis11
Flexible and Effective Object Tiering for Heterogeneous Memory Systems11
ODGS: Dependency-Aware Scheduling for High-Level Synthesis with Graph Neural Network and Reinforcement Learning11
Mitigating the Bandwidth Wall via Data-Streaming System–Accelerator Co-Design11
AG-SpTRSV: An Automatic Framework to Optimize Sparse Triangular Solve on GPUs11
MLKAPS: Machine Learning and Adaptive Sampling for HPC Kernel Auto-tuning11
Accelerating Nearest Neighbor Search in 3D Point Cloud Registration on GPUs11
Towards high scalability and fine-grained parallelism on distributed HPC platforms10
COX : Exposing CUDA Warp-level Functions to CPUs10
GraphSER: Distance-Aware Stream-Based Edge Repartition for Many-Core Systems10
Accelerating Parallel Structures in DNNs via Parallel Fusion and Operator Co-Optimization10
FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data Reuse10
Power Scheduling for Maximizing Throughput and Fairness in Co-running Applications10
Quantifying Resource Contention of Co-located Workloads with the System-level Entropy10
BridgeGC: An Efficient Cross-Level Garbage Collector for Big Data Frameworks9
NEM-GNN: DAC/ADC-less, Scalable, Reconfigurable, Graph and Sparsity-Aware Near-Memory Accelerator for Graph Neural Networks9
CLAP: Cross-Layer Adaptive Pipelining Inference Scheduling for Resource-Efficient Edge-Cloud Vision Systems9
EXPERTISE: An Effective Software-level Redundant Multithreading Scheme against Hardware Faults9
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions9
Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core Processor9
A Step toward Stateful HW-SW Migration: An Architecture-agnostic Checkpointing-rollback Toolchain9
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture9
A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks9
CML-PowF: Data Clustering Matching Based Low-overhead Multiple CPU Real-time Power Forecasting8
Efficient Cross-platform Multiplexing of Hardware Performance Counters via Adaptive Grouping8
Environmental Condition Aware Super-Resolution Acceleration Framework in Server-Client Hierarchies8
HEngine: A High Performance Optimization Framework on a GPU for Homomorphic Encryption8
A Decoupled Analytical Model for Tile Size Selection in Affine Programs8
HyGain: High-performance, Energy-efficient Hybrid Gain Cell-based Cache Hierarchy8
TPRepair: Tree-based Pipelined Repair in Clustered Storage Systems8
WSGraph: A Framework for Tackling Redundant and Irregular Data Access in Streaming Graph Processing8
RT-GNN: Accelerating Sparse Graph Neural Networks by Tensor-CUDA Kernel Fusion8
Towards High Performance QNNs via Distribution-Based CNOT Gate Reduction8
Multi-objective Hardware-aware Neural Architecture Search with Pareto Rank-preserving Surrogate Models8
PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power Constraints8
PowerMorph: QoS-Aware Server Power Reshaping for Data Center Regulation Service7
PyramidFFT: Rearchitecting FFT with Matrix-Aligned Nested-Radix for Hierarchical Scratchpad Memory on AI Accelerators7
RACER: Avoiding End-to-End Slowdowns in Accelerated Chip Multi-Processors7
RaNAS: Resource-Aware Neural Architecture Search for Edge Computing7
A Stable Idle Time Detection Platform for Real I/O Workloads7
SimTrace: Exploiting Spatial and Temporal Sampling for Large-Scale Performance Analysis7
Enabling Low-Latency, GPU-Efficient Serverless Inference with Model Swapping7
DTAP: Accelerating Strongly-Typed Programs with Data Type-Aware Hardware Prefetching7
Mobile-3DCNN: An Acceleration Framework for Ultra-Real-Time Execution of Large 3D CNNs on Mobile Devices7
Orchard: Heterogeneous Parallelism and Fine-grained Fusion for Complex Tree Traversals6
A Memory-Aware Sparse Matrix-Matrix Multiplication on Multicore Architectures6
MemoriaNova: Optimizing Memory-Aware Model Inference for Edge Computing6
Stripe-schedule Aware Repair in Erasure-coded Clusters with Heterogeneous Star Networks6
BLR-Krylov: A Single-GPU Iterative SpMM Framework with Communication Avoidance and Block Low-Rank Optimization6
Accelerating the Simulation of Parallel Workloads using Loop-Bounded Checkpoints6
Towards Optimizing Learned Index for High Performance, Memory Efficiency and NUMA Awareness6
EDAS: Enabling Fast Data Loading for GPU Serverless Computing6
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing6
OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUs6
Exploring Data Layout for Sparse Tensor Times Dense Matrix on GPUs6
Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing Environments6
Toward Comprehensive Design Space Exploration on Heterogeneous Multi-core Processors6
Pac-PIM: A Parallel Communication Framework for Commodity Processing-in-memory Systems6
PaTGen: Temporal Similarity-Driven Proxy Benchmark Generation Method for Cloud Workloads6
HiSo: Co-optimizing the Intra-layer and Inter-layer Scheduling Schemes with the Hybrid Data Flow for PIM Architectures6
MoonPoly : Bridging Code Generation and Adaptive Execution via Micro-Kernel Polymerization for Optimizing Dynamic-Shape Tensor Operators6
Lightweight Code Outlining for Android Applications6
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations6
x Meta : SSD-HDD-hybrid Optimization for Metadata Maintenance of Cloud-scale Object Storage6
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography6
CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph Processing5
GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing5
Improving Utilization of Dataflow Unit for Multi-Batch Processing5
RaKV: A Write-Optimized LSM Store for Cloud Block Storage with Robust SLA5
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation5
Koala: Efficient Pipeline Training through Automated Schedule Searching on Domain-Specific Language5
LLMSGHD: Co-Designing Large Language Model Software and Hardware for Efficient Inference5
Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUs5
JiuJITsu: Removing Gadgets with Safe Register Allocation for JIT Code Generation5
Capability-Based Efficient Data Transmission Mechanism for Serverless Computing5
Architecting Optically Controlled Phase Change Memory5
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory5
WIPE: A Write-Optimized Learned Index for Persistent Memory5
Efficient and Scalable Hybrid Parallelization of Unstructured Computational Fluid Dynamics with Geometric Multigrid5
Shift-CIM: In-SRAM Alignment To Support General-Purpose Bit-level Sparsity Exploration in SRAM Multiplication5
MetaEC: An Efficient and Resilient Erasure-Coded KV Store on Disaggregated Memory5
CoolDC: A Cost-Effective Immersion-Cooled Datacenter with Workload-Aware Temperature Scaling5
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers4
SAL: Optimizing the Dataflow of Spin-based Architectures for Lightweight Neural Networks4
MicroProf : Code-level Attribution of Unnecessary Data Transfer in Microservice Applications4
Optimizing OpenCL Barrier Synchronization and Memory Efficiency on Multi-Core DSPs4
Diamond Tiling for Periodic Stencil Loop Nests by Means of Transitive Closure-based Extraction of Overlapping Iteration Spaces4
Consequence-based Clustered Architecture4
Scale-out Systolic Arrays4
BLG-Tuning: Benchmark-Based Low-Cost General-Purpose I/O Modeling and Tuning4
Efficient Flexible Edge Inference for Mixed-Precision Quantized DNN using Customized RISC-V Core4
Fragment: Efficient DNN Checkpoint with Relaxed Model Consistency4
High Performance Singular Value Decomposition on GPU Architectures4
FlowPix: Accelerating Image Processing Pipelines on an FPGA Overlay using a Domain Specific Compiler4
CoNST: Code Generator for Sparse Tensor Networks4
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs4
An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPU4
Tetrahedrangel: An Energy-Efficient Coalescing Temporal Prefetcher4
Architectural Support for Sharing, Isolating and Virtualizing FPGA Resources4
SplitZNS: Towards an Efficient LSM-Tree on Zoned Namespace SSDs4
TSN Cache: Exploiting Data Localities in Graph Computing Applications4
Abakus: Accelerating k -mer Counting with Storage Technology4
GNΩSIS: Lessons Learned in Generating a High-Level Synthesis Dataset4
HeapKV: Enabling Efficient Garbage Collection for KV-Separated LSM Stores on Modern SSDs4
MLIR Dialects for DAE-based Modeling Languages4
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator4
Protocol and Software Co-optimization: Overcoming the PCIe Transfer Bottleneck in dGPU Accelerated In-kernel ML4
Address/Data Instruction Steering in Clustered General Purpose Processors4
Analytical Modeling of Set-Associative Caches for Optimizing Tensor Operations4
Rethinking Variable-Length Encoding: Exploiting Bit Sparsity for Parallel Decoding in LLM Accelerators4
GraphService: Topology-aware Constructor for Large-scale Graph Applications3
Optimization of Sparse Matrix Computation for Algebraic Multigrid on GPUs3
SSD-SGD: Communication Sparsification for Distributed Deep Learning Training3
FORTIFY: Feature-Oriented Representation and Graph Topology Integration for Path-Level Vulnerability Detection3
A Low-latency On-chip Cache Hierarchy for Load-to-use Stall Reduction in GPUs3
The Impact of Page Size and Microarchitecture on Instruction Address Translation Overhead3
Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance3
Data Deduplication Based on Content Locality of Transactions to Enhance Blockchain Scalability3
STen: Productive and Efficient Sparsity in PyTorch3
Heterogeneous Confidential Computing System for Large Language Models: A Survey3
Context-Aware Inlining: Using Call-Stack Profiles for Fast and Smaller Binaries3
Towards Minimized Cost for Wide Stripe Generation via Incremental Stripe Merging in Distributed Storage3
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel3
Supporting Dynamic Program Sizes in Deep Learning-Based Cost Models for Code Optimization3
CPU-GPU Workload Distribution during Throughput-Oriented LLM Inference on Single-GPU Systems3
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks3
FastCC: A System-Algorithm Co-Design for Connected Components Computation on Large Power-Law Graphs3
Corrigendum: Unified and Efficient Factor Graph Accelerator Design for Robotic Optimization3
GenCNN: A Partition-Aware Multi-Objective Mapping Framework for CNN Accelerators Based on Genetic Algorithm3
NPUMeter: Automatic Operator Optimization for Ascend NPU with Accurate Analytical Performance Models3
An Optimized GPU Implementation for GIST Descriptor3
High-performance Deterministic Concurrency Using Lingua Franca3
Jointly Optimizing Job Assignment and Resource Partitioning for Improving System Throughput in Cloud Datacenters3
Towards Efficient Extendible Perfect Hashing for Hybrid PM-DRAM Memory3
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access3
PARALiA: A Performance Aware Runtime for Auto-tuning Linear Algebra on Heterogeneous Systems3
A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep Learning3
Bubble-Swap Flow Control3
Parallax: Performance Prediction for Training–Inference Co-Execution3
PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow Architectures3
Puppeteer: A Random Forest Based Manager for Hardware Prefetchers Across the Memory Hierarchy3
MatXtract: Sparsity-Aware Matrix Transformation via Cascaded Compute Density EXtraction for SpMV3
PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware Computation3
SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical Workloads3
CheriMore: On-Demand Vertical Memory Expansion for Capability Serverless Runtime3
AdaptiveKV: Accelerating KV Cache Offloading with a Bandwidth-Adaptive Memory Allocation Mechanism3
Compressing and Accelerating Sparse CNNs Using Sign-Reserved Toeplitz Filters and Input Activation Density-aware Dataflow3
Cheetah: Accelerating Dynamic Graph Mining with Grouping Updates3
Making Root Cause Localization on FPGA Simulation Tools Robust3
LitTLS: Lightweight Thread-Level Speculation on Little Cores2
Winols: A Large-Tiling Sparse Winograd CNN Accelerator on FPGAs2
A Lock-free RDMA-friendly Index in CPU-parsimonious Environments2
Partitioned Scheduling and Analysis for a Typed DAG Task on Heterogeneous Multi-Cores2
ShieldCXL: A Practical Obliviousness Support with Sealed CXL Memory2
ZNSFQ: An Efficient and High-Performance Fair Queue Scheduling Scheme for ZNS SSDs2
Understanding Silent Data Corruption in Processors for Mitigating its Effects2
HAVIT: An Efficient H ardware- A ccelerator for V ision Tra2
Membrane: Accelerating Database Analytics with DRAM-Based PIM Filtering and Schema Denormalization2
UniTe: A Universal Tensor Abstraction for Capturing Spatial Relationships2
SAC: An Ultra-Efficient Spin-based Architecture for Compressed DNNs2
At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC Workloads2
Approx-RM: Reducing Energy on Heterogeneous Multicore Processors under Accuracy and Timing Constraints2
Unveiling and Evaluating Vulnerabilities in Branch Predictors via a Three-Step Modeling Methodology2
PiDRAM: A Holistic End-to-end FPGA-based Framework for Processing-in-DRAM2
Assessing the Impact of Compiler Optimizations on GPUs Reliability2
GOLDYLOC: Global Optimizations & Lightweight Dynamic Logic for Concurrency2
SuccinctKV: a CPU-efficient LSM-tree Based KV Store with Scan-based Compaction2
In-SRAM Parallel Data Shuffle2
AutoCO: An Online Continuous Optimization System for Phase-Changing Database Workloads.2
SPIRIT: Scalable and Persistent In-Memory Indices for Real-Time Search2
QuCloud+: A Holistic Qubit Mapping Scheme for Single/Multi-programming on 2D/3D NISQ Quantum Computers2
PARADISE: Criticality-Aware Instruction Reordering for Power Attack Resistance2
AccelHSA: Modeling Single-ISA Heterogeneous GPU Architectures2
ReSA: Reconfigurable Systolic Array for Multiple Tiny DNN Tensors2
Lock-Free High-performance Hashing for Persistent Memory via PM-aware Holistic Optimization2
Fast One-Sided RDMA-Based State Machine Replication for Disaggregated Memory2
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization2
ALOHA: Accelerating Leveled Fully Homomorphic Encryption with Cryptography-Specific Architectures2
ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads2
Cerberus: Triple Mode Acceleration of Sparse Matrix and Vector Multiplication2
ReIPE: Recycling Idle PEs in CNN Accelerator for Vulnerable Filters Soft-Error Detection2
NICE: Deep Neural Network Acceleration via Hardware-Friendly Index Assisted Compression2
HAIR: Halving the Area of the Integer Register File with Odd/Even Banking2
SpecTerminator: Blocking Speculative Side Channels Based on Instruction Classes on RISC-V2
Conflict Management in Vector Register Files2
Comperity: A Stateless Spiking Neural Network Accelerator with Differential Spike Encoding2
Compiler Support for Sparse Tensor Computations in MLIR2
WA-Zone: Wear-Aware Zone Management Optimization for LSM-Tree on ZNS SSDs1
gHyPart: GPU-friendly End-to-End Hypergraph Partitioner1
Second-level Caches: Not for Instructions1
DAG-Order: An Order-Based Dynamic DAG Scheduling for Real-Time Networks-on-Chip1
TEA+ : A Novel Temporal Graph Random Walk Engine with Hybrid Storage Architecture1
PRAGA: A Priority-Aware Hardware/Software Co-design for High-Throughput Graph Processing Acceleration1
Accurate Models of NVIDIA Tensor Cores1
Knowledge-Augmented Mutation-Based Bug Localization for Hardware Design Code1
An Optimized Framework for Matrix Factorization on the New Sunway Many-core Platform1
Rethinking the Energy Efficiency of SNNs and ANNs: A Perspective from Neural Network Design1
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer1
Energy-efficient In-Memory Address Calculation1
Dynamic Power Management Through Multi-agent Deep Reinforcement Learning for Heterogeneous Systems1
RegCPython: A Register-based Python Interpreter for Better Performance1
Reinforcement Learning on Data-Dependence Graphs for Custom Instruction Identification1
Characterizing and Optimizing LDPC Performance on 3D NAND Flash Memories1
Critical Data Backup with Hybrid Flash-Based Consumer Devices1
SpMARD: A Sparse-Sparse Matrix Multiplication Accelerator with Reconfigurable Dataflow for DNN Workloads1
Attack and Defense: Enhancing Robustness of Binary Hyper-Dimensional Computing1
Maximizing Data and Hardware Reuse for HLS with Early-Stage Symbolic Partitioning1
Phronesis: Efficient Performance Modeling for High-dimensional Configuration Tuning1
CoMeT: An Integrated Interval Thermal Simulation Toolchain for 2D, 2.5D, and 3D Processor-Memory Systems1
DELTA: Memory-Efficient Training via Dynamic Fine-Grained Recomputation and Swapping1
Fixed-point Encoding and Architecture Exploration for Residue Number Systems1
A 2 : Towards Accelerator Level Parallelism for Autonomous Micromobility Systems1
The Design of an Efficient Lossy Compressor for Time Series Databases1
Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUs1
Unleashing Parallelism with Elastic-Barriers1
AOBO: A Fast-Switching Online Binary Optimizer on AArch641
JUNO++: Optimizing ANNS and Enabling Efficient Sparse Attention in LLM via Ray Tracing Core1
SEED: Speculative Security Metadata Updates for Low-Latency Secure Memory1
Cppless: Single-Source and High-Performance Serverless Programming in C++1
An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing1
WindowQuant: Mixed-Precision KV Cache Quantization Based on Window-Level Similarity for VLMs Inference Optimization1
ASSG: Enhanced Workload Balancing via Adaptive State Scheduling Granularity Approach for Stateful Distributed Stream Processing1
SLAP: Segmented Reuse-Time-Label Based Admission Policy for Content Delivery Network Caching1
0.077749967575073