Quantity

For each crystal and each symmetry-inequivalent site, the constant-volume monovacancy formation energy

E_f = E_vac − (N − 1)/N · E_bulk

where E_bulk is the perfect N-atom supercell energy and E_vac the (N−1)-atom supercell energy, both after ion-only relaxation at fixed cell. One row per inequivalent site; DB key (structure_id, potential, site_index).

Corpus

The campaign ran the full 834 elemental structures (89 elements) — the same all-polymorph corpus as the elastic track — and the site ships every one of them. The leaderboard defaults to a ground-state facet (toggle: Ground-state only / All polymorphs / Metastable), because the high-energy polymorphs stress the relaxers and the consensus alike. The ground-state subset is 88 structures / 166 inequivalent sites / 88 elements.

Archive snapshot (vacancy_archive_20260627, benchmark commit 8e047ba): 233,388 done, 24,142 unconverged, 1,821 error rows across 60 potentials.

Construction & sites

Relaxation protocol

Stage Cell DoF Optimiser
Bulk pre-relax cell + ions BFGS
Perfect supercell (E_bulk) ions only, fixed cell FIRE
Vacancy supercell (E_vac) ions only, fixed cell FIRE

fmax = 1e-2 eV/Å, max 1000 steps. A row is done only when all three feeding relaxations reach fmax; unconverged rows are labelled and excluded from the spread. Failures record the last 4000 characters of the traceback; sticky CUDA-context faults trigger a worker respawn.

Reference & metrics

There is no DFT/MP comparand. The reference is the cross-potential consensus of E_f per site (median / MAD / min / max / n), plus each potential's signed deviation, overlaid where available with a curated experimental value. The live leaderboard columns:

Metric What
vacancy_consensus_mae (headline) median |E_f − per-site cross-potential median E_f| (eV), medianAbsSkipNull, lower better
vacancy_consensus_mbe signed median deviation from the per-site consensus (eV)
vacancy_lit_mae median |E_f − experimental E_f| (eV) over the curated overlay (sparse, curated metals)
convergence % per-potential lookup column, done / (done + unconverged)
runtime / GPU per-potential cost panel (avg_runtime_ms, avg_gpu_mem_mb)

The headline vacancy_consensus_mae is the implemented MAE-vs-median — the cross-potential spread is the reference model, so it ranks "most typical", not "most correct" (read it alongside the experiment column). The per-site spread (n_potentials, median_e_f, mad_e_f, min_e_f, max_e_f) feeds the structures drill-in; the rollup's n_converged counts converged potentials at a site, not "sites a potential converged".

Experimental literature overlay

A curated table of 20 common elemental metals in src/assets/elemental/vacancy/vacancy_literature.csv, sourced from the Korhonen–Puska–Nieminen compilation (Phys. Rev. B 51, 9526 (1995)) plus primary positron-annihilation / dilatometry papers (Fluss–Smedskjaer, Triftshäuser, Maier–Seeger, Feder–Charbnau, …). The overlay is phase-gated: a value is attached to an element's ground-state site only when the corpus 0 K ground state IS the experimentally-measured phase (spacegroup match). For 5 metals (Ag, Mg, In, Na, Ti) the lowest-hull DFT polymorph differs from the measured phase, so they are listed in the CSV but not attached — leaving 15 phase-matched, single-site overlay elements (Al, Cu, Au, Ni, Pt, Pd, Fe, W, Mo, Ta, Nb, Cr, V, Pb, Co). Finding: the consensus systematically under-predicts noble-metal and Ni vacancies vs experiment (Pt −0.65, Pd −0.51, Au −0.46, Cu −0.20, Ni −0.30 eV) while matching Al / Fe / Mo / Nb / Ta within ~0.1 eV; a few early-3d metals (Cr +0.84, Co +0.43) instead over-predict.

Experimental references

Every experimental value, its measured phase, and its primary source (the machine-readable table is vacancy_literature.csv, downloadable and on R2). A marks a value listed but not attached to the leaderboard, because the corpus 0 K ground state is a different polymorph from the experimentally-measured phase.

Element E_fexp (eV) Phase On site Primary source
Al 0.67 fcc Fluss–Smedskjaer 1978; Ehrhart LB III/25
Cu 1.28 fcc Korhonen–Puska–Nieminen 1995; Triftshäuser–McGervey
Au 0.90 fcc Korhonen–Puska–Nieminen 1995; Triftshäuser–McGervey
Ni 1.79 fcc Korhonen–Puska–Nieminen 1995
Pt 1.35 fcc Korhonen–Puska–Nieminen 1995
Pd 1.70 fcc Korhonen–Puska–Nieminen 1995; Ehrhart LB III/25
Pb 0.58 fcc Feder–Nowick 1967; Ehrhart LB III/25
Fe 2.00 bcc De Schepper 1983; Ehrhart LB III/25
W 3.60 bcc Korhonen–Puska–Nieminen 1995; Maier–Seeger 1979
Mo 3.00 bcc Korhonen–Puska–Nieminen 1995
Ta 2.80 bcc Korhonen–Puska–Nieminen 1995
Nb 2.60 bcc Korhonen–Puska–Nieminen 1995
Cr 2.00 bcc Korhonen–Puska–Nieminen 1995; Ehrhart LB III/25
V 2.20 bcc Korhonen–Puska–Nieminen 1995
Co 1.34 hcp Maier 1979
Ag 1.11 fcc Korhonen–Puska–Nieminen 1995; Triftshäuser–McGervey
Mg 0.79 hcp Tzanetakis 1976; Ehrhart LB III/25
In 0.48 bct Suzuki–Nagai et al. 2001 (PRB 63, 180101)
Na 0.42 bcc Feder–Charbnau 1966
Ti 1.55 hcp Ehrhart LB III/25; specific-heat (Shestopal)

Bibliography

Published-DFT reference (PBE)

A second overlay compares each potential to published PBE-DFT vacancy formation energies from Medasani, Haranczyk, Canning & Asta, Comp. Mater. Sci. 101, 96 (2015) (doi:10.1016/j.commatsci.2015.01.018). That study uses the identical constant-volume definition E_f = E_vac − (N−1)/N·E_bulk with relaxed atomic positions at fixed volume, so it is a like-for-like comparand. We take the uncorrected PBE Ef column (not the surface-energy-corrected ˜Ef) — PBE because the benchmarked MLIPs are PBE surrogates — phase-gated to the corpus ground state exactly like the experimental overlay, giving 24 phase-matched metals (vacancy_dft_reference.csv):

Phase Metals (PBE E_f, eV)
fcc Al 0.65 · Cu 1.09 · Au 0.41 · Ni 1.46 · Pt 0.74 · Pd 1.21 · Rh 1.74 · Ir 1.62 · Ca 1.18
bcc Fe 2.20 · W 3.31 · Mo 2.74 · Ta 2.82 · Nb 2.77 · Cr 2.77 · V 2.27
hcp Co 1.96 · Hf 2.24 · Os 3.03 · Re 3.40 · Ru 2.71 · Sc 1.86 · Tc 2.79 · Zn 0.42

This is the most direct accuracy measure: the cross-potential consensus reproduces published PBE DFT to within ~0.1 eV on most metals (Al −0.02, Cu −0.01, Au +0.03, Ir 0.00, Fe −0.11, Co −0.20), and the best OAM/OMAT-trained models reach a vs-PBE-DFT MAE of ~0.05 eV. It also explains the experiment gap: PBE itself under-predicts noble-metal vacancies (PBE Au 0.41 vs exp 0.90; PBE Pt 0.74 vs exp 1.35), so the MLIPs' apparent under-prediction vs experiment is largely inherited from PBE, not an MLIP failure — the models are faithful PBE surrogates.

Robustness — why everything is a median

~4 % of raw E_f are non-physical blow-ups (up to ~1e25 eV) from potentials that "converge" to a garbage PES minimum, so all metrics are median / median-of-absolute. The ground-state-default view is clean (median E_f range −1.76 … 6.92 eV; only the Br₂ / I₂ molecular crystals have an unphysical negative consensus median, with a few open non-metal sites — P, B — carrying comparably large MAD). For ~200 hard metastable sites >50 % of potentials blow up, so even the per-site median is unreliable — hence the ground-state default and the "most typical, not most correct" caveat.

Status: complete. The pending stress-test polymorphs (zero-symmetry P1 cells with ~100 sites × ~800-atom supercells) were purged on 2026-06-27; existing results were kept and the rest left NaN (no zero-filling). 797 structures have a result for all 60 potentials (the clean, rectangular set); 37 are ragged (NaN for ≥1 potential).

Provenance & verification (deep, code-cited) → /methodology/deep/vacancy

See the leaderboard → · ← Methodology overview

Code

Symmetry-resolved monovacancy formation energy for one crystal (vacancy_compute.py in mlip-elemental-benchmark):

from ase.build import bulk
from vacancy_compute import compute_vacancies
calc = ...   # any ASE calculator for your MLIP (e.g. mace_mp(model="medium", device="cuda"))

res = compute_vacancies(bulk("Fe", "bcc", a=2.83, cubic=True), calc)
for s in res["sites"]:
    print(s["wyckoff_symbol"], s["formation_energy"])   # per-inequivalent-site E_f (eV)