Skip to content

Perov-5: what the runs show

Seven training runs of the alignment model, three without the curriculum and three with materials held out give more than one number per model. This page reads them for what they say about training, about failures and about the shared space, and turns each result into something you can use. Every number is generated from findings.json; the runs are described on Alignment across seeds.

How the model learns · Where it fails · What the shared space contains · In practice

How the model learns

The curriculum builds the structures first

With the curriculum, the contrastive weight starts at zero and grows over the first 1,500 epochs. In a paired run (seed 6, the same starting weights and data order), structure matching reaches 88.3 % at epoch 50 with the curriculum and 8.5 % without it; without the curriculum it is still 36.2 % at epoch 300. The alignment behaves the other way round: without the curriculum the cosine is already 0.88 at epoch 10, against 0.03 with it.

Structure matching (most likely element), seed 6

Structure matching during the first 300 epochs, with and without the curriculum01002003000255075epochstructure matching (%)with the curriculumwithout

Cosine between the two latents, seed 6

Cosine during the first 300 epochs, with and without the curriculum010020030000.20.40.60.8epochcosinewith the curriculumwithout

After the full training, over three and seven seeds:

cosine L2 structure matching, sampled element
with the curriculum (7 seeds, 2,200 epochs) 0.954 ± 0.011 0.301 ± 0.034 86.1 ± 5.0 %
without (3 seeds, 2,500 epochs) 0.949 ± 0.009 0.316 ± 0.028 67.1 ± 7.6 %

What this means for you

The curriculum does not change how well the modalities align in the end. It decides whether the model also learns to rebuild crystals. Keep training.contrastive_warmup_epochs on when you train on your own data.

Reconstruction peaks early, then swings

In the seed-6 run, structure matching is highest at epoch 300 (93.7 %), when the cosine is only 0.82. From there to the end it moves between 77.8 % and 93.7 % from one snapshot to the next, and ends at 88.5 % (cosine 0.94).

Structure matching and cosine over 2,200 epochs, seed 6 (29 snapshots)

Structure matching and cosine over the whole training, seed 605001,0001,5002,0000255075epochstructure matching (%)cosine × 100

What this means for you

The last epoch is not the best checkpoint for reconstruction. Save checkpoints during training and choose one on a validation measure. This is one seed; the swings may differ for others.

Early alignment does not predict the final one

At epoch 50 the cosine of the seven seeds ranges from 0.17 to 0.73; at the end from 0.94 to 0.96. The order of the seeds at epoch 100 has a rank correlation of only 0.32 with their final order (figure on Alignment across seeds).

What this means for you

Do not judge a model, or stop a run, on its alignment after a few dozen epochs. A short training, such as the 30 epochs of the hosted Studio, gives a rough model whose quality depends on the seed.

Alignment and reconstruction are independent

Across the seven seeds, the correlation between the final cosine and structure matching is 0.17 (sampled element) and 0.02 (most likely element). The seed with the highest cosine (0.965) has the lowest structure matching (78.2 %).

What this means for you

Report both. A high cosine says nothing about how well crystals are rebuilt.

Where it fails

Failures are seed noise, not hard materials

Of the 18,928 materials, 99.94 % are reconstructed by at least one of the seven seeds and 51.9 % by all seven; only 11 are missed by every seed. Whether one seed fails on a material says almost nothing about another seed (correlation 0.06).

Number of materials reconstructed by k of the seven seeds

Materials reconstructed by k of the seven seeds7 of 7 seeds9,8316 of 7 seeds6,2695 of 7 seeds2,0424 of 7 seeds5123 of 7 seeds1652 of 7 seeds691 of 7 seeds290 of 7 seeds11

Right element per site, seven seeds pooled (A and B: cations; X: anions)

Share of materials with the right element on each sitesite X299.96 %site X199.90 %site X399.89 %site A96.65 %site B95.47 %

So several seeds can be combined. Adding the element probabilities of the seeds, site by site, and taking the most likely element gives the exact composition for 99.77 % of the materials with three seeds and 99.99 % with seven, against 92.1 % for one seed on average (85.6–96.2 %).

What this means for you

Three models trained with different seeds, with their element probabilities added, remove almost all composition errors on the training materials. It costs three trainings and no change to the model.

The bottleneck is naming the cation

79 % of the failed reconstructions have a wrong element, and almost always on a cation site: the A site is right for 96.7 % of the materials and the B site for 95.5 %, the three anion sites for 99.9 % or more. When the composition is right, 97.8 % of the crystals match; the lattice length is off by 0.10 Å on average.

On unseen materials the whole loss is the cation

trained on 80 %, three seeds A site right B site right anion sites right match when the composition is right lattice error
materials seen in training 97.4 % 97.2 % 99.9 % 96.2 % 0.104 Å
held-out materials 86.5 % 84.0 % 99.4 % 96.0 % 0.115 Å

What this means for you

The geometry generalises: lattice and positions are as good on unseen materials as on seen ones. What does not generalise is choosing the A and B elements. Work on the model should go there first; candidates should have their composition checked first.

Take the most likely element

The training script's evaluation samples the element of each site. Taking the most likely element instead raises structure matching by 3.9 points on average (2.8 to 5.4 over the seven seeds).

The decoder knows when it is unsure

Take the lowest of the five top probabilities of a decoded crystal (one per site) as its confidence. It needs no reference structure.

confidence share of materials (seen) reconstructed correctly (seen) share (held-out) reconstructed correctly (held-out)
≥ 0.9 77.9 % 96.0 % 63.7 % 82.3 %
0.6 to 0.9 17.6 % 74.5 % 26.9 % 51.5 %
< 0.6 4.5 % 49.0 % 9.4 % 34.3 %
all 100 % 90.1 % 100 % 69.5 %

Area under the ROC curve: 0.81 on seen materials (seven seeds), 0.75 on held-out materials (three seeds).

What this means for you

Rank or filter decoded crystals by this confidence before spending computing time on them.

Chemistry matters little, some cations more

Structure matching by anion set, seven seeds pooled:

anions O₂N ONF O₂F O₃ ON₂ O₂S N₃
structure matching 91.7 % 91.2 % 91.0 % 91.0 % 89.3 % 89.1 % 87.2 %

The cations with the lowest structure matching (elements with at least 100 materials):

site lowest highest
A Os 82 %, Cs 84 %, Ir 85 %, Rb 85 %, B 85 %, K 87 % Sn 95 %, Bi 94 %, Sb 94 %, Ga 94 %
B Ca 82 %, Hf 85 %, Mn 86 %, B 86 %, Ta 86 %, La 87 % Cu 94 %, Ni 94 %, Sb 93 %, Ga 93 %

What the shared space contains

The modalities are aligned before the projection heads

The contrastive loss of this model acts on the normalised encoder outputs. There the cosine is 0.954 ± 0.011. After the projection heads, which feed the decoder, the same pairs have a cosine of -0.978 ± 0.002: the two projected vectors point in nearly opposite directions, and the decoder reads their mean.

In the benchmark

The representation task compares the two latents in the space where each model aligns them, declared per method. The cosine in both spaces is recorded on each method's page.

The 128-number latents use about two directions

The participation ratio, the effective number of directions a set of vectors uses, is 2.2 ± 0.0 for the structure latents and 2.3 ± 0.1 for the property latents, out of 128; the joint latent that the decoder reads uses 8.5 ± 0.8.

What this means for you

Two scalar properties can organise two directions, and the alignment pulls the structure latent onto them. A richer second modality, such as a diffraction pattern or a density of states, is the way to a richer shared space (roadmap).

The property modality of Perov-5 is thin

96.1 % of the materials (18,193) have a direct band gap of 0, and the 18,928 materials share only 903 distinct pairs of (formation enthalpy, band gap). 94 % of the materials have the same pair as at least ten others; the largest group has 261.

Retrieval works in one direction

top 1 median rank
from a structure, find its properties 0.230 186
from the properties, find the structure 0.003 228

Ranks are among all 18,928 materials; a material sharing its properties with others is not counted against. From properties to a structure, many crystals are equally right answers.

What this means for you

This is why inverse design in MEIDNet is a search and not a lookup: a material family fixes the sites, rules remove impossible compositions, and the search moves through the latent space (How MEIDNet works).

Less stable materials align better and reconstruct worse

By fifths of the formation enthalpy, from the lowest to the highest:

formation enthalpy up to (eV/atom) 0.88 1.22 1.54 2.00 5.16
cosine 0.952 0.948 0.950 0.955 0.963
structure matching 91.6 % 91.6 % 91.2 % 89.9 % 86.0 %

The correlation between a material's cosine and its formation enthalpy is 0.47 ± 0.06 over the seven seeds.

What a structure alone predicts

On the 3,786 held-out materials, with the property of the five nearest training materials in the structure latent (three seeds):

formation enthalpy (eV/atom) direct band gap (eV)
five nearest neighbours in the structure latent 0.060 ± 0.004 0.087 ± 0.003
nearest property latent (cross-modal) 0.062 ± 0.004 0.131 ± 0.016
linear model on the structure latent 0.144 ± 0.015 0.152 ± 0.023
mean of the training materials 0.563 0.167

Mean absolute errors. The band-gap numbers are small because most gaps are zero.

What this means for you

The properties decoded from the joint latent have an error of only 0.008 eV/atom, but the joint latent contains the true properties: that is a reconstruction. The error to expect for a new structure is the first row, about 0.06 eV/atom.

In practice

What the results above suggest for anyone training or using such a model:

  1. Keep the curriculum. It is what teaches the model to rebuild crystals.
  2. Train three seeds, not one. Results differ between seeds, and their errors are independent.
  3. Choose the checkpoint on a validation measure, not at the last epoch.
  4. Decode with the most likely element, and add the probabilities of the seeds when you have several.
  5. Rank the decoded crystals by confidence (the lowest top probability over the sites) and check the composition first.
  6. Confirm with an MLIP, then DFT. The guide runs the first of these steps.

Steps 2, 4 and 5 are measured here on reconstruction, with the alignment model. They are not yet options of meidnet generate.