We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:

References & Citations


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Distributed, Parallel, and Cluster Computing

Title: Scalable communication for high-order stencil computations using CUDA-aware MPI

Abstract: Modern compute nodes in high-performance computing provide a tremendous level of parallelism and processing power. However, as arithmetic performance has been observed to increase at a faster rate relative to memory and network bandwidths, optimizing data movement has become critical for achieving strong scaling in many communication-heavy applications. This performance gap has been further accentuated with the introduction of graphics processing units, which can provide by multiple factors higher throughput in data-parallel tasks than central processing units. In this work, we explore the computational aspects of iterative stencil loops and implement a generic communication scheme using CUDA-aware MPI, which we use to accelerate magnetohydrodynamics simulations based on high-order finite differences and third-order Runge-Kutta integration. We put particular focus on improving intra-node locality of workloads. In comparison to a theoretical performance model, our implementation exhibits strong scaling from one to $64$ devices at $50\%$--$87\%$ efficiency in sixth-order stencil computations when the problem domain consists of $256^3$--$1024^3$ cells.
Comments: 17 pages, 15 figures
Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Computational Physics (physics.comp-ph); Fluid Dynamics (physics.flu-dyn)
Cite as: arXiv:2103.01597 [cs.DC]
  (or arXiv:2103.01597v1 [cs.DC] for this version)

Submission history

From: Johannes Pekkilä [view email]
[v1] Tue, 2 Mar 2021 09:44:42 GMT (32kb)

Link back to: arXiv, form interface, contact.