Method & caveats

For every (potential × Materials-Project grain boundary) the campaign rebuilds the CSL boundary with pymatgen at the model's own equilibrium lattice constant (≈ 15 Å per grain, in-plane orthogonalised for the 60°/120° coincidence cells) and relaxes it, scoring the relaxed grain-boundary energy γ (J/m²) against the Materials-Project DFT reference for the same boundary (from the Zheng et al. 2020 set, matched to each MP parent). So — unlike the GB energetics track — this is a real accuracy benchmark, not a cross-model consensus.

Three build protocols, chosen with the tab above the leaderboard. Shifted scans the in-plane rigid-body shift — the γ-surface — on a static grid to find the minimum-energy registry before relaxing; zero relaxes the naive [0,0] registry with no scan; hybrid (the default) uses the shifted result for twist boundaries and the zero-shift result for tilt boundaries. Across the fleet the hybrid protocol is the most accurate — γ MAE 0.1105 J/m² vs 0.1258 (shifted) and 0.1263 (zero) — and wins for 51 of 59 potentials; the most accurate single model is equiformer_v3_oam at 0.0447 J/m². The static γ-surface scan's benefit over the naive build (≈ 14 %) largely survives the subsequent relaxation, which is why the hybrid — keeping the scan only where it pays off — edges out either pure protocol.

Two metric modes. The default scores γ vs DFT. The excess-volume mode scores each model's grain-boundary excess volume against the DFT excess volume, on the paired ev_quality=='reference' subset with |ev| ≤ 5 Å (blow-ups excluded); it reports EV MAE / bias (Å) and an EV-vs-DFT correlation. The per-boundary spread plot stays γ-based in both modes.

The headline is MAE vs DFT (J/m² for γ, Å for excess volume); RMSE exposes the heavy σ7 [111] tail, and Pearson r the rank fidelity. Converged is each potential's boundaries that produced a γ out of 327; a handful of unstable potentials (orb_v2, eqV2, matris, some MatBench checkpoints) crash the relaxer on many cells and sit at the bottom on both accuracy and convergence — read the two together. On the periodic-table map, an element rendered in amber was attempted but produced no converged γ (every potential errored on it — mostly the lanthanides, where the hexagonal-GB build hits a cell-convention failure); it is clickable, and its boundaries are listed below with the dominant exception. Blank cells are simply not in the benchmark. The γ-outlier filter (leaderboard controls) drops any (potential × GB) point whose relaxed γ falls outside a chosen [min, max] window before the MAE/RMSE are aggregated — useful for asking "how accurate is the field once the handful of unphysical blow-ups (up to ~140 J/m²) are set aside?"; it is off by default so the headline stays the raw accuracy. The cubic (fcc/bcc) and hexagonal (hcp) families are built by different geometry paths and carry different error scales; the crystal-family toggle separates them. The grid is 59 potentials × 327 grain boundaries, and the leaderboard n is the paired count scored under the active protocol/metric. Full protocol — the pymatgen build, the γ-surface scan and the DFT matching: Pure GB methodology →.