Page 1 of 1

Large energy change upon rigid translation of H2O in a finite-temperature RPAR calculation

Posted: Thu Oct 01, 2026 4:37 am
by sc_han

Dear VASP developers,

I observe a large change in the finite-temperature RPA energy when rigidly translating an isolated H2O molecule within the same cell.

I compared two structures with identical cell and internal geometry, with all three atoms shifted by +0.20 Å along z. Throughout this post, Δ denotes the energy of the shifted structure minus that of the reference structure, i.e.,

Code: Select all

ΔFHF       = FHF(shifted)       - FHF(reference)
ΔEc(RPA)   = Ec(RPA)(shifted)   - Ec(RPA)(reference)
ΔE(HF+RPA) = E(HF+RPA)(shifted) - E(HF+RPA)(reference)

For ENCUT = ENCUTGW = 280 eV, NBANDS = 54, and the same 25-point k-point mesh, I obtained:

Code: Select all

                         ΔFHF         ΔEc(RPA)     ΔE(HF+RPA)
PRECFOCK = Fast        -153.315 meV   -10.593 meV   -163.908 meV

The corresponding DFT-only calculations give the same free energy and forces to the printed precision. The translation dependence therefore appears in the HF+RPA calculation and is dominated by the change in FHF.

Using a 50-point k-point mesh with the same NBANDS = 54 gave a similar shifted-minus-reference difference:

Code: Select all

ΔE(HF+RPA) = -164.974 meV

A separate PRECFOCK = Normal pair gave:

Code: Select all

ΔFHF       =  +9.792 meV
ΔEc(RPA)   =  -0.190 meV
ΔE(HF+RPA) =  +9.603 meV

In a follow-up calculation with PRECFOCK = Normal, the translation-dependent energy difference was substantially smaller (9.6 meV, compared with 163.9 meV for Fast).

I also noticed a substantial difference in the estimated memory requirement. At the same NTAUPAR = 9, the estimated maximum memory per rank was 54,501 MB for Normal, compared with 27,184 MB for Fast. This makes PRECFOCK = Normal difficult to use when the available memory per rank is limited.

Is there a recommended way to reduce the memory requirement of PRECFOCK = Normal while retaining its improved energy stability? In particular, I would appreciate any guidance on which parallelization or numerical settings are most effective for reducing the memory footprint in this case.

Thank you very much for your time and for any advice you can provide.

Relevant settings:

Code: Select all

ALGO = RPAR
LFINITE_TEMPERATURE = .TRUE.
NOMEGA = 16
ENCUT = ENCUTGW = 280 eV
ISMEAR = -1
SIGMA = 0.15 eV
PRECFOCK = Fast
NMAXFOCKAE = 1

Best regards,
Seungchang Han


Re: Large energy change upon rigid translation of H2O in a finite-temperature RPAR calculation

Posted: Thu Oct 01, 2026 8:25 am
by merzuk.kaltak

Dear Seungchang,

thank you for your report. It would be great if you could post your INCAR, KPOINTS, POSCARs and OUTCAR (maybe also the stdout) of the two runs.
Without those files I can only speculate what is going on here (for instance NBANDS=54 seems too low for a converged correlation energy, but sufficient for the exact exchange energy).

For accurate exact exchange energies we recommend PRECFOCK=Normal. However, this typically increases the memory requirement compared to the default settings. There are in general not many settings that one can use to tweak memory requirement for ACFDT/RPA calculations . There are basically two tags you can set.

  • ENCUTGW sets the number of plane-wave coefficients for the response function. Since the FFTs are done in the Born-von-Karman supercell of the k-point grid times the unit-cell, the code requires to store matrices of the size NKPTS x N(plane-waves) ~ NKPTS x E_{cut}^{3/2}.

  • NTAUPAR is used to distribute the imaginary time points that represent the time dependence of the response functions. The code stores NOMEGA response functions in the BvK supercell and distributes them into NTAUPAR groups. The least memory is used for NTAUPAR=1 because all MPI ranks share one matrix. Typically this tag is set automatically based on the available memory (read from the Linux kernel on each compute node) and the code picks the best value for the run based on the available resources.

If you still hit a wall in terms of memory, there is not much more you can do, except using more compute nodes (in addition to use less MPI ranks).