1. Thermodynamic integration (primary)

The headline route is calphy non-equilibrium thermodynamic integration (NETI): each phase's absolute Gibbs free energy is fixed by reversibly switching the Hamiltonian, over a finite alchemical path λ ∈ [0, 1], between the MLIP and an analytically solvable reference —

For each (potential, element) the campaign runs 10 single-temperature free-energy kernels — 5 solid + 5 liquid — on the reduced-temperature grid

T = [0.60, 0.75, 0.90, 1.00, 1.10] × Tm_exp

widened to [0.40 … 1.00] × Tm_exp for a handful of very soft, low-melting elements where the default grid sits too close to (or above) the true Tm to bracket a crossing. Cells are ~2000–2700 atoms; each kernel runs N_equil = 10,000 equilibration steps followed by N_switch = 25,000 switching steps, n_iter = 1 (no repeat averaging — see the switching-convergence study for why 25k is adequate).

Each phase's five (T, G) points are fit linearly in T; Tm is the temperature where the solid and liquid G(T) lines cross. Two quality gates run before the fit:

2. Switching-convergence study

Longer switching paths dissipate less irreversible work, so a natural question is whether the production N_switch = 25,000 setting is simply too fast to have converged. To check, Mo and Re were rerun at three switching lengths plus a Richardson extrapolation to N_switch → ∞:

Element Tm exp (K) 25k 50k 100k Richardson (N→∞)
Mo 2896 2507 2540 2577 2614
Re 3459 2703 2772 2800 2828

Quadrupling the switching time from 25k to 100k steps recovers only +70 K (Mo) and +97 K (Re) — a small fraction of the ~300–650 K gap to experiment — and extrapolating all the way to infinite switching time (zero dissipation) still leaves both elements ~10 % (Mo) / ~18 % (Re) low. Production therefore stays at N_switch = 25,000: the residual gap to experiment is real model-plus-method error (the MLIP's PES and/or the NETI free-energy protocol itself), not an artefact of insufficiently slow switching. This is the evidence behind the TI ensemble's systematic low bias reported in §4.

3. Coexistence (secondary)

The independent cross-check builds an explicit two-phase solid–liquid cell (~4000 atoms) and brackets the melting point directly, rather than via free energies:

Coexistence is run for a subset of (potential, element) pairs — where an explicit two-phase cell could be constructed and stably equilibrated — not the full elemental grid TI covers.

4. The two methods disagree by design

TI and coexistence are kept separate, not averaged, because they fail in opposite regimes and agreeing with both would only hide that:

On the trusted-4 clean ensemble (sign-change crossings only, no-sign-change extrapolations excluded), the TI headline is:

(The campaign report's ti_results_trusted.csv instead pools the no-sign-change extrapolations into its per-element means, giving MAE ≈ 24.1 %, MBE ≈ −23.1 %, 29 of 61 within ±20 % over that larger, mixed-quality pool — the regen tool's cross-check certifies against these pooled numbers.)

The honest reading is that a uniform ~20+ % low bias is a real, characterized systematic offset (confirmed by the switching-convergence study above), not noise to be averaged away. The more reassuring number sits alongside it: the inter-potential spread — how much the potentials agree with each other — is only ~16 %, tighter than the 21.6 % MAE against experiment. The models under-predict Tm in close agreement with one another; they are more precise than they are accurate. Where TI and coexistence both ran for the same (potential, element), treat the pair as two independent, differently-biased estimates of the same quantity — not as a single averaged number.

5. Data caveats

See the leaderboard → · ← Methodology overview