Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Dominik Ernst; Georg Hager; Jonas Thies; Gerhard Wellein

doi:10.1177/1094342020965661

Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Dominik Ernst, Georg Hager, Jonas Thies, Gerhard Wellein

Research output: Contribution to journal › Article › Scientific › peer-review

5 Citations (Scopus)

Abstract

General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall & skinny matrices, which are much taller than wide. NVIDIA’s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code generation combined with autotuning. For a large range of matrix sizes in the domain of interest we achieve at least 2/3 of the roofline performance and often substantially outperform state-of-the art CUBLAS results on an NVIDIA Volta GPGPU.

Original language	English
Pages (from-to)	5-19
Number of pages	15
Journal	International Journal of High Performance Computing Applications
Volume	35
Issue number	1
DOIs	https://doi.org/10.1177/1094342020965661
Publication status	Published - Jan 2021
Externally published	Yes

Keywords

CUDA
GPU
Performance engineering
complex
matrix multiplication
tall & skinny

Access to Document

10.1177/1094342020965661

Cite this

@article{a7383e42f5be4f7a8c9c4b91e7f59bdc,

title = "Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs",

abstract = "General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall & skinny matrices, which are much taller than wide. NVIDIA{\textquoteright}s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code generation combined with autotuning. For a large range of matrix sizes in the domain of interest we achieve at least 2/3 of the roofline performance and often substantially outperform state-of-the art CUBLAS results on an NVIDIA Volta GPGPU.",

keywords = "CUDA, GPU, Performance engineering, complex, matrix multiplication, tall & skinny",

author = "Dominik Ernst and Georg Hager and Jonas Thies and Gerhard Wellein",

year = "2021",

month = jan,

doi = "10.1177/1094342020965661",

language = "English",

volume = "35",

pages = "5--19",

journal = "International Journal of High Performance Computing Applications",

issn = "1094-3420",

publisher = "SAGE Publishing",

number = "1",

}

TY - JOUR

T1 - Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

AU - Ernst, Dominik

AU - Hager, Georg

AU - Thies, Jonas

AU - Wellein, Gerhard

PY - 2021/1

Y1 - 2021/1

N2 - General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall & skinny matrices, which are much taller than wide. NVIDIA’s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code generation combined with autotuning. For a large range of matrix sizes in the domain of interest we achieve at least 2/3 of the roofline performance and often substantially outperform state-of-the art CUBLAS results on an NVIDIA Volta GPGPU.

AB - General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall & skinny matrices, which are much taller than wide. NVIDIA’s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code generation combined with autotuning. For a large range of matrix sizes in the domain of interest we achieve at least 2/3 of the roofline performance and often substantially outperform state-of-the art CUBLAS results on an NVIDIA Volta GPGPU.

KW - CUDA

KW - GPU

KW - Performance engineering

KW - complex

KW - matrix multiplication

KW - tall & skinny

UR - http://www.scopus.com/inward/record.url?scp=85092373486&partnerID=8YFLogxK

U2 - 10.1177/1094342020965661

DO - 10.1177/1094342020965661

M3 - Article

SN - 1094-3420

VL - 35

SP - 5

EP - 19

JO - International Journal of High Performance Computing Applications

JF - International Journal of High Performance Computing Applications

IS - 1

ER -

Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this