GPU 6.6.0 deadlock in HSE06 KPOINTS_OPT band step (OpenMP/OpenACC collision, vhamil_trace_mu)
================================================================
1. SUMMARY
================================================================
In a PBE->HSE06 workflow running across GPU (A100) and CPU nodes, we found a reproducible deadlock in the GPU build during the HSE combined calculation, specifically when the job transitions from the main-KPOINTS SCF to the KPOINTS_OPT band-path (non-SCF) extension
within the same run. The SCF itself converges correctly; the hang occurs the instant KPOINTS_OPT starts.
perf profiling and a gdb backtrace point to an OpenMP host-thread barrier (in the NVIDIA HPC SDK OpenMP runtime, libnvomp.so) getting stuck while multiple OpenMP threads attempt to enter the same OpenACC-offloaded region inside vhamil_trace_mu() (the HSE Fock-exchange routine). This looks like an OMP_NUM_THREADS-driven issue rather than a KPAR/MPI-rank issue. Details and reproduction steps below.
================================================================
2. BUILD / ENVIRONMENT
================================================================
VASP build (GPU): vasp.6.6.0-gpu-skylake
VASP build (CPU): vasp.6.6.0; original binary compiled with -march=native (AVX-512) caused crashes on AMD nodes -> rebuilt with -march=znver3
Parallel libraries: OpenMPI 4.1.6, NVIDIA HPC SDK OpenMP runtime (libnvomp.so), HPC-SDK version: 26.5 (CUDA 13.2)
CPU math libraries: AMDBLIS / AMDLIBFLAME / AMDSCALAPACK / AMDFFTW (AOCL)
GPU nodes: Intel Xeon Platinum 8275CL (Cascade Lake) + 8x A100 per node
CPU nodes (band stage): AMD EPYC 7R13 (Zen3); AMD EPYC 9R14 (Zen4c, 128 physical cores)
Test material: CsRb2HoF6 (f-electron material, contains Ho)
k-points: Main KPOINTS: Gamma 6x6x6 -> 112 irreducible k-points (SCF)
KPOINTS_OPT: line-mode band path, 280 k-points, processed in batches of 112 (non-SCF extension)
================================================================
3. STEPS TO REPRODUCE
================================================================
1. HSE06 INCAR (LHFCALC=.TRUE., etc.) together with a main KPOINTS file (112 k-points, for SCF) and a KPOINTS_OPT file (280 k-points, line-mode band path) in the same job, run with the GPU build.
2. One MPI rank per GPU, KPAR = <number of GPUs> (NUMA-aware placement via an OpenMPI rankfile), OMP_NUM_THREADS left at its default (6 in our environment).
3. SCF proceeds normally to convergence.
4. Immediately after SCF convergence, the job hangs at the "Start KPOINTS_OPT" line — reproduced identically at KPAR=2 and KPAR=3 (KPAR=3 verified by restarting from an already-converged WAVECAR).
Log excerpt at the hang point:
Start KPOINTS_OPT (optional k-point list driver)
k-point batch 1-112\280 (no further updates to vasp.out/OUTCAR)
At this point the process is alive and consumes >250% CPU, but strace shows only repeated futex(FUTEX_WAKE)/sched_yield() calls — a busy-spin pattern with no actual computation.
================================================================
4. DIAGNOSTIC EVIDENCE
================================================================
- HSE SCF on the main KPOINTS (112 k-points) converges cleanly and reproducibly (F = -62.073214 eV, 22 CGA iterations, ~25.9h). GPU utilization during this stage is ~95% (KPAR parallelism confirmed working correctly). The hang occurs strictly after this, at the KPOINTS_OPT transition.
- perf profiling (perf record --call-graph dwarf) while the job is hung shows ~40-41% of all sampled stack frames sitting inside hxiExecuteHostTreeBarrierWithTasks in libnvomp.so (NVIDIA HPC SDK OpenMP runtime). This is an OpenMP host-side thread-tree barrier, not an MPI communication routine.
- KPAR=2 and KPAR=3 hang in the exact same function at nearly identical sample ratios (40.44% vs 41.04%). This argues against KPAR (MPI rank count) as the trigger and points instead toward OMP_NUM_THREADS (OpenMP threads per rank).
- A follow-up gdb backtrace pinpoints the hang inside vhamil_trace_mu() (the HSE Fock-exchange computation routine), at an "!$ACC LOOP VECTOR" directive — an OpenACC GPU-offload entry point. This is consistent with multiple OpenMP host threads simultaneously attempting to enter the same GPU-offload region and colliding.
- Setting OMP_NUM_THREADS=1 (KPAR unchanged) eliminates the deadlock, confirming the OpenMP-thread hypothesis, but the loss of thread-level parallelism makes this impractical: the run could not finish a single k-point within 20 hours.
================================================================
5. CURRENT WORKAROUND AND ITS LIMITS
================================================================
We split the combined stage into a GPU SCF-only stage and a CPU-only KPOINTS_OPT (band) stage: SCF converges on GPU with KPAR=2 (validated, ~26h/material), and the converged WAVECAR is handed off to a CPU-only restart (rebuilt AMD znver3 binary, NCORE=1, KPAR=2) for the band extension. Under this path, an np=50 CPU run showed genuine progress with no deadlock (gdb instruction-pointer movement over ~3h15m), but no run has yet completed through to EIGENVAL/PROCAR_OPT — this is not yet a fully validated workaround.
Two secondary notes from this workaround, possibly useful for reproducibility:
(a) VASP rounds NBANDS up to a multiple of the per-KPAR-group rank count on WAVECAR restart. If this count does not evenly divide NBANDS=50 (e.g. 32 total ranks / KPAR=2 -> 16 ranks/group forces NBANDS up to 64), the extra bands (51-64) appear to be randomly initialized.
Choosing a per-group rank count that divides 50 evenly (10 or 25) avoided this — see question 4 below.
(b) Separately, OpenMPI 4.1.6's rankfile mapper fails silently (no error message) when physical and hyperthread cores are mixed and the total rank count reaches 41 or more. This may be unrelated to VASP itself but is noted here in case it affects reproduction on other systems.
================================================================
6. QUESTIONS
================================================================
1) Is it a known issue that the GPU build can deadlock in HSE06 + KPOINTS_OPT runs when multiple OpenMP threads (OMP_NUM_THREADS >= 2) concurrently enter an OpenACC-offloaded region (e.g. inside vhamil_trace_mu)? Is there a relevant patch, or a version where this was addressed?
2) Is there a supported way to apply different OMP_NUM_THREADS values to the SCF stage versus the KPOINTS_OPT extension stage — either by switching thread count mid-job, or must these always run as separate jobs?
3) Is running the HSE06 band-structure step (KPOINTS_OPT) CPU-only, without GPU offload, an officially recommended/supported pattern?
Or is there a different recommended configuration (build flags, environment variables) to make this step run stably on GPU?
4) When restarting from a WAVECAR, is the automatic rounding of NBANDS up to a multiple of the per-KPAR-group rank count intended behavior, and is there documentation confirming the added bands are initialized safely (reproducibly) in that case?
5) The VASP wiki's "GPU ports of VASP" page notes a past bug in NVIDIA HPC-SDK 22.1/22.2 where the OpenACC version could not be combined with OpenMP-threading, fixed in HPC-SDK 22.3+. Our build uses HPC-SDK 26.5 (CUDA 13.2). Given the symptom overlap (OpenMP-thread / OpenACC-region interaction), is it possible this is a related regression in a different HPC-SDK version?
Attached: INCAR/KPOINTS/KPOINTS_OPT/POSCAR, a trimmed OUTCAR up to the hang point, the GPU and CPU makefile.include files, the exact rankfile and launch command used, and a text summary of the perf/gdb findings above (the original perf.data and gdb backtrace files from that diagnostic session were not preserved, so we're summarizing the numbers here instead). We're happy to reproduce and attach the raw perf.data / gdb backtrace on request.
