VASPACC 6.5.0 issue with Rocky 10.2 and latest nvhpc
We have upgraded our cluster to Rocky 10.2, which included bringing all the major system software and utility components to their latest stable releases.
I've built for Icelake nodes with 4x A30 GPUs each. The difference between the arch/xxx and my makefile.include is below. The build goes cleanly. The build is done with the commands:
module purge; module load nvhpc compiler-rt tbb mkl and then make DEPS=1 -j 16 all
The issue is that a user whose configuration ran fine in Rocky 8, now crashes on GPU memory consumption.
I've been dragged through the mod with Google AI as to what is going on. It's made several suggestions, based on a supposed difference between how things work with the latest SLURM, Rocky 10, and nvhpc. The suggestions bounce between the INCAR file settings and the script to run a job; some of these are quite essoteric and it's hard to believe that anyone could have figured it out. So, my basic questions are:
1) What changes are to be expected for the INCAR and sbatch files to get the GPUs to be allocated, one per process, not overloading each other onto the same GPU?
2) Is it known that 6.5.0 won't work at all? If so, I'll have my users provide me a 6.6.1 version. But so far, I've not found anything claiming a major issue for 6.5.0.
Other stacks run fine on the GPUs (RStudio/R, JupyterLab/Python, ML codes, etc). So, it seems the machine is well-configured... except for the apps configuration for VASPACC.
BTW, the user is using vasp_std for his runs.
Thank you for any hints you can pass my way.
Stephen
[easybuild@ebicelake vasp.6.5.0]$ diff arch/makefile.include.nvhpc_ompi_mkl_omp_acc makefile.include
20,22c20,26
< CC = mpicc -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp
< FC = mpif90 -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp
< FCL = mpif90 -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp -c++libs
---
> #CC = mpicc -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp
> #FC = mpif90 -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp
> #FCL = mpif90 -acc -gpu=cc60,cc70,cc80,cuda11.8 -mp -c++libs
> CC = mpicc -acc -gpu=cc75,cc80 -mp
> FC = mpif90 -acc -gpu=cc75,cc80 -mp
> FCL = mpif90 -acc -gpu=cc75,cc80 -mp -c++libs
>
82c86
< MKLLIBS = -lmkl_intel_lp64 -lmkl_pgi_thread -lmkl_core -pgf90libs -mp -lpthread -lm -ldl
---
> MKLLIBS = -lmkl_intel_lp64 -lmkl_gnu_thread -lmkl_core -pgf90libs -mp -lpthread -lm -ldl