CUDA-DTM: Distributed transactional memory for GPU clusters
- Samuel Irving,
- Sui Chen,
- Lu Peng(corresponding author),
- ,
- Maurice Herlihy,
- Christopher J. Michael
- Louisiana State University,
- Brown University
Related Event
Title
Event type
ConferenceDate
06/19/2019 - 06/21/2019Location
Abstract
We present CUDA-DTM, the first ever Distributed Transactional Memory framework written in CUDA for large scale GPU clusters. Transactional Memory has become an attractive auto-coherence scheme for GPU applications with irregular memory access patterns due to its ability to avoid serializing threads while still maintaining programmability. We extend GPU Software Transactional Memory to allow threads across many GPUs to access a coherent distributed shared memory space and propose a scheme for GPU-to-GPU communication using CUDA-Aware MPI. The performance of CUDA-DTM is evaluated using a suite of seven irregular memory access benchmarks with varying degrees of compute intensity, contention, and node-to-node communication frequency. Using a cluster of 256 devices, our experiments show that GPU clusters using CUDA-DTM can be up to 115x faster than CPU clusters.
Publication Information
Output type
Original language
English (US)Pages from-to (Number of pages)
Pages 183-199 (17 pages)Publication milestones
- Published - 2019
Publication status
Publisher
SpringerPublication series
- Publication series name: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
ISSN (Print): 0302-9743
ISSN (Electronic): 1611-3349
Volume: 11704 LNCS
ISBN (Print)
9783030312763Publication IDs
- Scopus: 85075595144
Host publication title
Networked Systems - 7th International Conference, NETYS 2019, Revised Selected PapersHost publication editors
- Mohamed Faouzi Atig
- Alexander A. Schwarzmann
