VASP 6.5.1: recurrent "internal error in rot.F / EDWAV: the gradient is not orthogonal"
Dear VASP developers,
I repeatedly hit the internal error below in a 287–297-atom DFT+U perovskite supercell (VASP 6.5.1). The error: internal error in: rot.F (...) EDWAV: internal error, the gradient is not orthogonal (...). It has killed 32 jobs so far (about 30% of cases) of SQS generated bulk supercells. It occurs either after 1 to 56 successful ionic steps. A rerun from the crashed job's CONTCAR with identical inputs gets past the crash point most of the times. It happens both with cold starts (ISTART=0, ICHARG=2: 13 crashes) or with warm-started relaxations (ISTART=1, ICHARG=0: 19 crashes), both with Gamma-point calculations (but with vasp_std bin - I have not tried with vasp_gam) and with 4 irreducible KPOINTS.
I have encountered a (maybe) related anomaly in some of the jobs that crashed: a wrong electron count on the FINAL SCF iteration of an ionic step. In relaxations of this cell I see occasional number of electron lines off NELECT by 0.003–0.12 e. Every instance checked is on the last SCF iteration of an ionic step, i.e. the iteration whose density gives the forces. But some of the cases that showed this anomaly eventually do converged after restoring the proper electron count.
What I have tried (None of these removed the crash):
NELMDL (0 and −5); three mixing/algorithm combinations (ALGO=Normal and ALGO=All, default and conservative AMIX/BMIX/AMIX_MAG BMIX_MAG); LREAL=.FALSE. (failed earlier); LDAU=.FALSE.; ISPIN=1. I also ruled out NBANDS exhaustion (ample headroom).
I have not tried: IWAVPR changes, ALGO=Damped/Fast for the relaxations, ADDGRID, LMAXMIX=6, or a different MPI/compiler build. I would rather ask you before changing anything that could affect results across an ongoing campaign.
System and settings (identical in all crashed jobs)
- Sr0.95Ti0.3Fe0.6Cu0.1O3−x perovskite, 3×4×5 supercell (60 B sites: 18 Ti, 36 Fe, 6 Cu; 57 Sr; 170–180 O; 287–297 atoms);
fixed lattice about 11.76 × 15.72 × 19.54 Å . - PBE+U, Dudarev (LDAUTYPE=2): U(Ti)=4.91, U(Fe)=4.5, U(Cu)=6.47 eV. Collinear AFM A-type order.
- POTCARs: PAW_PBE O, Ti_sv, Fe, Cu, Sr_sv.
- ENCUT=600, PREC=Accurate, LREAL=Auto, ISYM=0, ISMEAR=0, SIGMA=0.05, NBANDS=1344, ALGO=All (IALGO 58), ISEARCH=1,
AMIX=0.05, BMIX=0.01, AMIX_MAG=0.1, BMIX_MAG=0.01, LASPH=.TRUE., LMAXMIX=4, EDIFF=1E-6, NELM=900,
IBRION=2 (POTIM 0.5), EDIFFG=−0.03. - Parallelisation: NCORE=8, 64 MPI ranks per node, 1 thread per rank, KPAR=1 (Γ) or 2/4 (k-mesh).
Build:
- vasp.6.5.1.
- Compiler and libraries:
- GNU Fortran 11.5.0 (Red Hat 11.5.0-2) via Intel MPI's
mpif90; - -O2 -march=native;
- MKL 2021.4.0 (gf_lp64, sequential, ScaLAPACK, BLACS-IntelMPI);
- CPP options: -DMPI -DMPI_BLOCK=8000 -Duse_collective -DscaLAPACK -DCACHE_SIZE=4000 -Davoidalloc -Dvasp6 -Dtbdyn
-Dfock_dblbuf -DVASP_HDF5 -DVASP2WANNIER90. - makefile.include is attached.
- The binary was built against Intel MPI 2021.16 (library path embedded in the binary).
- At run time the environment loads oneapi/mpi/latest, which has pointed to Intel MPI 2021.18.1
- GNU Fortran 11.5.0 (Red Hat 11.5.0-2) via Intel MPI's
- The binary contains no AVX-512 instructions. It was built with -march=native on an AMD EPYC 7262
I attach two bundles with complete inputs and outputs for 7 crashed jobs and 2 successful comparison jobs, the full crash list, and the build and hardware details.
Questions:
- Is this a known issue in 6.5.1 for ALGO=All with large spin-polarised DFT+U cells? Is there a fix or patch? Is the occasional wrong electron count on the final SCF iteration a known symptom of the same problem?
- In the relaxation case, is there a recommended setting that avoids the failure without changing the converged result? For example IWAVPR, ALGO=Damped, or a smaller trial step.
- Could the build/run MPI mismatch (built 2021.16, run 2021.18.1) or
-march=nativebuilt on Zen 2 and run on Zen 4 plausibly cause this? Should we rebuild?
Best regards,
[MS edit: changed a few lines that were misinterpreted as LaTeX]