Skip to content

Perov-5

18,928 cubic ABX3 perovskites with relaxed structures, formation enthalpy (heat_all, eV/atom) and direct band gap (dir_gap, eV). The CDVAE split of the Castelli et al. (2012) perovskite set; the dataset behind the published MEIDNet model and the live Studio.

18,928 materials · split: CDVAE: 11,356 train / 3,787 val / 3,785 test · properties: heat_all, dir_gap · dataset source

perov5-v1.1protocol
3tasks
5models
5baselines
54candidates per method

Inverse design · Property prediction · Representation · Alignment across seeds · What the runs show · Choose a model and generate · Protocol · Reported in the paper · Submit a method

Inverse design

Propose cubic ABX3 perovskites for three band-gap targets in three chemical families (54 candidates); score them for stability with an MLIP, uniqueness, novelty and, where a DFT value exists, the target band gap. Export: CSV · JSON.

#MethodSUN ↑Stable ↑Unique ↑Novel ↑ΔHf ↓DFT hit ↑DFT knownDelivered ↑ParamsTraining dataAddedLinks
1Encoder screeningbaseline0.8520.8521.0001.000-1.28–0540.70 MPerov-5, all 18,928 materials2026-10-03
2Random samplingbaseline0.7780.8891.0000.889-1.340.000654––2026-10-03
3MEIDNet (first early-fusion model)0.6480.9441.0000.704-1.290.18816540.70 MPerov-5, all 18,928 materials2026-10-03
4MEIDNet (published model)0.6110.9621.0000.660-1.520.05618530.70 MPerov-5, all 18,928 materials2026-10-03
5MEIDNet (alignment training, seed 3)0.6110.8721.0000.830-1.320.2508470.70 MPerov-5, all 18,928 materials2026-10-05
6MEIDNet (shorter training)0.5560.9231.0000.654-1.500.05618520.70 MPerov-5, all 18,928 materials2026-10-03
How to read the inverse design table

Rows are sorted by the first metric column; click any header to sort by another. Colours show the rank within each column, from dark (last) to yellow (first); ↑ marks metrics where higher is better, ↓ where lower is better. Hover a header for its definition.

  • SUN: stable, unique and novel candidates divided by the budget of 54; a candidate that was not delivered counts as a failure.
  • Stable: fraction of delivered candidates whose MACE-MP-0 formation energy after relaxation is at most 0.10 eV/atom, against elemental reference phases (the criterion of meidnet screen and of the paper).
  • Unique: fraction of delivered candidates whose composition does not repeat an earlier one.
  • Novel: fraction of delivered candidates whose composition is not in the data set (training, validation and test splits).
  • ΔHf (eV/atom): median MACE-MP-0 formation energy of the delivered candidates.
  • DFT hit: among candidates whose A, B and X sites match a Perov-5 entry (so their DFT band gap is known), the fraction within 0.5 eV of the target.
  • DFT known: candidates with a known DFT band gap (the denominator of DFT hit).
  • Delivered: candidates delivered out of the budget of 54.
  • Valid: fraction of the budget that passes every rule of the family. Listed on each method's page.
  • SUN count: stable, unique and novel candidates. Listed on each method's page.
  • Budget: candidates requested. Listed on each method's page.
  • Generated: candidates generated (reported in the paper). Listed on each method's page.

SUN measures stability, uniqueness and novelty, not whether a candidate has the target band gap: methods that never return a composition of the data set, such as the two baselines, set a high reference for it. Of the compositions that pass the family's rules, 40 of the 106 oxides are in the data set and no chalcogenide or halide is, so novelty separates methods among oxides only. The DFT check is the only measure of the target and is possible for oxides whose A and B sites match a Perov-5 entry (40 of the 106 rule-passing oxides; 26 of them have a non-zero direct gap).

SUN by chemical family. Stable, unique and novel candidates out of the budget of each family; hover a cell for the number of novel candidates.

Methodoxidechalcogenidehalide
Encoder screening18/1811/1817/18
Random sampling12/1814/1816/18
MEIDNet (first early-fusion model)2/1815/1818/18
MEIDNet (published model)0/1816/1817/18
MEIDNet (alignment training, seed 3)10/1812/1811/18
MEIDNet (shorter training)0/1814/1816/18

Property prediction

Predict formation enthalpy and direct band gap from the crystal structure of the 3,785 test materials. Export: CSV · JSON.

#MethodMAE ΔH ↓RMSE ΔH ↓R² ΔH ↑MAE gap ↓MAE gap > 0 ↓ParamsTraining dataAddedLinks
1MEIDNet (published model)0.4040.5390.4752.7921.6760.70 MPerov-5, all 18,928 materials†2026-10-03
2Composition k-NNbaseline0.5110.6740.1790.1632.044–Perov-5 train (11,356)2026-10-03
3MEIDNet (first early-fusion model)0.5220.6980.11913.46911.9980.70 MPerov-5, all 18,928 materials†2026-10-03
4Training meanbaseline0.5660.744-0.0000.1732.135–Perov-5 train (11,356)2026-10-03
5MEIDNet (shorter training)1.0091.391-2.4953.2921.9780.70 MPerov-5, all 18,928 materials†2026-10-03

† Trained on all materials of the data set, the test split included: for these rows the scores are measured on materials the model has seen, and are not comparable with rows trained on the training split only.

How to read the property prediction table

Rows are sorted by the first metric column; click any header to sort by another. Colours show the rank within each column, from dark (last) to yellow (first); ↑ marks metrics where higher is better, ↓ where lower is better. Hover a header for its definition.

  • MAE ΔH (eV/atom): mean absolute error of the formation enthalpy (heat_all).
  • RMSE ΔH (eV/atom): root-mean-square error of the formation enthalpy.
  • R² ΔH: coefficient of determination of the formation enthalpy (0 = no better than the mean).
  • MAE gap (eV): mean absolute error of the direct band gap over all test materials; 96% of them have a gap of 0 eV.
  • MAE gap > 0 (eV): mean absolute error of the direct band gap on the test materials with a non-zero gap, the range that inverse design targets.
  • RMSE gap (eV): root-mean-square error of the direct band gap. Listed on each method's page.
  • R² gap: coefficient of determination of the direct band gap; not informative here, because the gap is zero for most materials. Listed on each method's page.
  • MAE ΔH ≠ 0 (eV/atom): mean absolute error of the formation enthalpy on materials where it is not zero. Listed on each method's page.
  • n: test materials evaluated. Listed on each method's page.

MEIDNet (published model): formation enthalpy, predicted vs. DFT

MEIDNet (published model): formation enthalpy (eV/atom), predicted vs. DFT (3,785 test materials)024024DFT formation enthalpy (eV/atom)predicted formation enthalpy (eV/atom)

MEIDNet (published model): direct band gap, predicted vs. DFT

MEIDNet (published model): direct band gap (eV), predicted vs. DFT (3,785 test materials)02460246DFT direct band gap (eV)predicted direct band gap (eV)

Composition k-NN: formation enthalpy, predicted vs. DFT

Composition k-NN: formation enthalpy (eV/atom), predicted vs. DFT (3,785 test materials)024024DFT formation enthalpy (eV/atom)predicted formation enthalpy (eV/atom)

Representation

How well the shared space holds both modalities on the test split: cross-modal retrieval, agreement of the two latents, and a k-nearest-neighbour probe of the structure latent. Export: CSV · JSON.

#MethodR@1 ↑R@5 ↑cos ↑k-NN MAE ΔH ↓k-NN MAE gap ↓ParamsTraining dataAddedLinks
1MEIDNet (published model)0.3360.8960.7720.0210.0550.70 MPerov-5, all 18,928 materials†2026-10-03
2MEIDNet (alignment training, 7 seeds)0.297 ± 0.0620.855 ± 0.0650.954 ± 0.0110.025 ± 0.0040.049 ± 0.0160.70 MPerov-5, all 18,928 materials†2026-10-05
3MEIDNet (shorter training)0.2970.8480.4800.0240.0470.70 MPerov-5, all 18,928 materials†2026-10-03
4MEIDNet (first early-fusion model)0.0280.1310.6550.0510.0640.70 MPerov-5, all 18,928 materials†2026-10-03
5Chance levelbaseline0.0030.0140.000––––2026-10-03
6Composition k-NNbaseline–––0.5110.163–Perov-5 train (11,356)2026-10-03

† Trained on all materials of the data set, the test split included: for these rows the scores are measured on materials the model has seen, and are not comparable with rows trained on the training split only.

How to read the representation table

Rows are sorted by the first metric column; click any header to sort by another. Colours show the rank within each column, from dark (last) to yellow (first); ↑ marks metrics where higher is better, ↓ where lower is better. Hover a header for its definition.

  • R@1: fraction of test materials for which, among the distinct property profiles of the test split, their own profile's latent is the nearest to their structure latent (materials with identical property values share one profile).
  • R@5: the same within the five nearest profiles.
  • cos: mean cosine similarity between the structure latent and the property latent of the same material, in the space where the model aligns them (before or after its projection heads; stated on the method's page).
  • k-NN MAE ΔH (eV/atom): formation-enthalpy error of a 5-nearest-neighbour probe: each test material takes the mean property of its five nearest training materials in the representation.
  • k-NN MAE gap (eV): direct-band-gap error of the same probe.
  • L2: mean L2 distance between the two latents of the same material (unit latents). Listed on each method's page.
  • cos, encoder outputs: the matched cosine between the normalised encoder outputs, before the projection heads. Listed on each method's page.
  • cos, projection heads: the matched cosine between the outputs of the projection heads. Listed on each method's page.
  • profiles: distinct property profiles among the test materials: the candidates of retrieval (chance level of R@1 is one over this number). Listed on each method's page.
  • n: test materials evaluated. Listed on each method's page.

Protocol

Version perov5-v1.1. Data: data/perov5 (meidnet download-data), split CDVAE: 11,356 train / 3,787 val / 3,785 test.

Changes of the protocol
  • perov5-v1.1 (2026-10-05). Representation: the two latents are compared in the space where the model aligns them (before or after its projection heads), declared per method; the cosine in both spaces is recorded. Inverse design: novelty is measured against all 18,928 materials of the data set instead of the training split. Records may give the mean and the standard deviation over several training runs. Rows trained on all materials are marked, because their test scores are not held-out.
  • perov5-v1 (2026-10-03). First version: three tasks, CDVAE split, novelty against the training split, latents compared after the projection heads.

Inverse design

Propose cubic ABX3 perovskites for three band-gap targets in three chemical families (54 candidates); score them for stability with an MLIP, uniqueness, novelty and, where a DFT value exists, the target band gap.

setting value
family perovskite_abx3, variants oxide, chalcogenide, halide
targets dir_gap 1.5, heat_all -0.1; dir_gap 2.5, heat_all -0.1; dir_gap 3.5, heat_all -0.1
budget 6 candidates per target and variant (54 in total), each passing the family's rules
MEIDNet search the generation settings of examples/perov5/meidnet.yaml, at most 6 rounds per target
stability MACE-MP-0 (medium), float32; BFGS with cell relaxation (FrechetCellFilter), fmax 0.05 eV/Å, at most 500 steps; formation energy against elemental phases at most 0.1 eV/atom
novelty composition absent from the data set (all 18,928 materials)
DFT check Perov-5 entries with the same A, B and X sites; hit within 0.5 eV of the target band gap

Property prediction

Predict formation enthalpy and direct band gap from the crystal structure of the 3,785 test materials.

Split: test. Reference implementation: meidnet.benchmark.

Representation

How well the shared space holds both modalities on the test split: cross-modal retrieval, agreement of the two latents, and a k-nearest-neighbour probe of the structure latent.

Split: test. Reference implementation: meidnet.benchmark.

Run the protocol for a MEIDNet checkpoint, or score the outputs of any other model:

python scripts/benchmarks.py run perov5 --method meidnet-2k            # a built-in method, every task
python scripts/benchmarks.py score perov5 --task property_prediction --predictions preds.csv
python scripts/benchmarks.py score perov5 --task inverse_design --candidates candidates.csv

preds.csv has one row per test material: id (the dataset's material_id) and pred_<property> for each property. candidates.csv has id, variant, target (1, 2, 3), formula, elements (JSON of the sites, e.g. {"A": "Ba", "B": "Ti", "X": "O"}) and cif (path of the structure). Stability needs MACE: set --mlip-python or MEIDNET_MLIP_PYTHON if it lives in another environment.

Reported in the paper

Numbers quoted from the paper use the paper's own settings (other targets, budget and splits), so they are listed here rather than ranked above.

record task values status
Candidates shipped with the paper, screened here Inverse design SUN 0.654, Stable 1.000, Unique 0.923, Novel 0.692, ΔHf -1.27, Delivered 26, SUN count 17 Computed here
MEIDNet (paper): inverse design Inverse design Generated 140, SUN count 19, SUN 0.136 Reported in the paper