Ensuring an effective user experience when managing and running scientific HPC software
2015 IEEE/ACM 1st International Workshop on Software Engineering for High Performance Computing in Science (SE4HPCS), pp. 56-59 (2015)
@inproceedings{cohen-2015b,
title = {Ensuring an effective user experience when managing and running scientific HPC software},
booktitle = {2015 IEEE/ACM 1st International Workshop on Software Engineering for High Performance Computing in Science (SE4HPCS)},
author = {Cohen, J. and Moxey, D. and Cantwell, C. D. and Austing, P. and Darlington, J. and Sherwin, S. J.},
year = {2015},
pages = {56-59},
url = {https://davidmoxey.uk/assets/pubs/2015-se4hpcs.pdf},
doi = {10.1109/SE4HPCS.2015.16},
abstract = {With CPU clock speeds stagnating over the last few years, ongoing advances in computing power and capabilities are being supported through increasing multi- and many-core parallelism. The resulting cost of locally maintaining large-scale computing infrastructure, combined with the need to perform increasingly large simulations, is leading to the wider use of alternative models of accessing infrastructure, such as the use of Infrastructure-as-a-Service (IaaS) cloud platforms. The diversity of platforms and the methods of interacting with them can make using them with complex scientific HPC codes difficult for users. In this position paper, we discuss our approaches to tackling these challenges on heterogeneous resources. As an example of the application of these approaches we use Nekkloud, our web-based interface for simplifying job specification and deployment of the Nektar++ high-order finite element HPC code. We also present results from a recent Nekkloud evaluation workshop undertaken with a group of Nektar++ users.}
}
Running a large simulation increasingly means using infrastructure maintained by someone else, and the variety of platforms and ways of reaching them makes that awkward for scientists. This position paper discusses approaches to the problem on heterogeneous resources, using Nekkloud, our web-based interface for specifying and deploying Nektar++ jobs, as the worked example, and reports on an evaluation workshop held with Nektar++ users.
Abstract
With CPU clock speeds stagnating over the last few years, ongoing advances in computing power and capabilities are being supported through increasing multi- and many-core parallelism. The resulting cost of locally maintaining large-scale computing infrastructure, combined with the need to perform increasingly large simulations, is leading to the wider use of alternative models of accessing infrastructure, such as the use of Infrastructure-as-a-Service (IaaS) cloud platforms. The diversity of platforms and the methods of interacting with them can make using them with complex scientific HPC codes difficult for users. In this position paper, we discuss our approaches to tackling these challenges on heterogeneous resources. As an example of the application of these approaches we use Nekkloud, our web-based interface for simplifying job specification and deployment of the Nektar++ high-order finite element HPC code. We also present results from a recent Nekkloud evaluation workshop undertaken with a group of Nektar++ users.