Perov-5: what the runs show¶
Seven training runs of the alignment model, three without the curriculum and three with materials held out give more than one number per model. This page reads them for what they say about training, about failures and about the shared space, and turns each result into something you can use. Every number is generated from findings.json; the runs are described on Alignment across seeds.
How the model learns¶
The curriculum builds the structures first¶
With the curriculum, the contrastive weight starts at zero and grows over the first 1,500 epochs. In a paired run (seed 6, the same starting weights and data order), structure matching reaches 88.3 % at epoch 50 with the curriculum and 8.5 % without it; without the curriculum it is still 36.2 % at epoch 300. The alignment behaves the other way round: without the curriculum the cosine is already 0.88 at epoch 10, against 0.03 with it.
Structure matching (most likely element), seed 6
Cosine between the two latents, seed 6
After the full training, over three and seven seeds:
| cosine | L2 | structure matching, sampled element | |
|---|---|---|---|
| with the curriculum (7 seeds, 2,200 epochs) | 0.954 ± 0.011 | 0.301 ± 0.034 | 86.1 ± 5.0 % |
| without (3 seeds, 2,500 epochs) | 0.949 ± 0.009 | 0.316 ± 0.028 | 67.1 ± 7.6 % |
What this means for you
The curriculum does not change how well the modalities align in the end. It decides whether the model also learns to rebuild crystals. Keep training.contrastive_warmup_epochs on when you train on your own data.
Reconstruction peaks early, then swings¶
In the seed-6 run, structure matching is highest at epoch 300 (93.7 %), when the cosine is only 0.82. From there to the end it moves between 77.8 % and 93.7 % from one snapshot to the next, and ends at 88.5 % (cosine 0.94).
Structure matching and cosine over 2,200 epochs, seed 6 (29 snapshots)
What this means for you
The last epoch is not the best checkpoint for reconstruction. Save checkpoints during training and choose one on a validation measure. This is one seed; the swings may differ for others.
Early alignment does not predict the final one¶
At epoch 50 the cosine of the seven seeds ranges from 0.17 to 0.73; at the end from 0.94 to 0.96. The order of the seeds at epoch 100 has a rank correlation of only 0.32 with their final order (figure on Alignment across seeds).
What this means for you
Do not judge a model, or stop a run, on its alignment after a few dozen epochs. A short training, such as the 30 epochs of the hosted Studio, gives a rough model whose quality depends on the seed.
Alignment and reconstruction are independent¶
Across the seven seeds, the correlation between the final cosine and structure matching is 0.17 (sampled element) and 0.02 (most likely element). The seed with the highest cosine (0.965) has the lowest structure matching (78.2 %).
What this means for you
Report both. A high cosine says nothing about how well crystals are rebuilt.
Where it fails¶
Failures are seed noise, not hard materials¶
Of the 18,928 materials, 99.94 % are reconstructed by at least one of the seven seeds and 51.9 % by all seven; only 11 are missed by every seed. Whether one seed fails on a material says almost nothing about another seed (correlation 0.06).
Number of materials reconstructed by k of the seven seeds
Right element per site, seven seeds pooled (A and B: cations; X: anions)
So several seeds can be combined. Adding the element probabilities of the seeds, site by site, and taking the most likely element gives the exact composition for 99.77 % of the materials with three seeds and 99.99 % with seven, against 92.1 % for one seed on average (85.6–96.2 %).
What this means for you
Three models trained with different seeds, with their element probabilities added, remove almost all composition errors on the training materials. It costs three trainings and no change to the model.
The bottleneck is naming the cation¶
79 % of the failed reconstructions have a wrong element, and almost always on a cation site: the A site is right for 96.7 % of the materials and the B site for 95.5 %, the three anion sites for 99.9 % or more. When the composition is right, 97.8 % of the crystals match; the lattice length is off by 0.10 Å on average.
On unseen materials the whole loss is the cation¶
| trained on 80 %, three seeds | A site right | B site right | anion sites right | match when the composition is right | lattice error |
|---|---|---|---|---|---|
| materials seen in training | 97.4 % | 97.2 % | 99.9 % | 96.2 % | 0.104 Å |
| held-out materials | 86.5 % | 84.0 % | 99.4 % | 96.0 % | 0.115 Å |
What this means for you
The geometry generalises: lattice and positions are as good on unseen materials as on seen ones. What does not generalise is choosing the A and B elements. Work on the model should go there first; candidates should have their composition checked first.
Take the most likely element¶
The training script's evaluation samples the element of each site. Taking the most likely element instead raises structure matching by 3.9 points on average (2.8 to 5.4 over the seven seeds).
The decoder knows when it is unsure¶
Take the lowest of the five top probabilities of a decoded crystal (one per site) as its confidence. It needs no reference structure.
| confidence | share of materials (seen) | reconstructed correctly (seen) | share (held-out) | reconstructed correctly (held-out) |
|---|---|---|---|---|
| ≥ 0.9 | 77.9 % | 96.0 % | 63.7 % | 82.3 % |
| 0.6 to 0.9 | 17.6 % | 74.5 % | 26.9 % | 51.5 % |
| < 0.6 | 4.5 % | 49.0 % | 9.4 % | 34.3 % |
| all | 100 % | 90.1 % | 100 % | 69.5 % |
Area under the ROC curve: 0.81 on seen materials (seven seeds), 0.75 on held-out materials (three seeds).
What this means for you
Rank or filter decoded crystals by this confidence before spending computing time on them.
Chemistry matters little, some cations more¶
Structure matching by anion set, seven seeds pooled:
| anions | O₂N | ONF | O₂F | O₃ | ON₂ | O₂S | N₃ |
|---|---|---|---|---|---|---|---|
| structure matching | 91.7 % | 91.2 % | 91.0 % | 91.0 % | 89.3 % | 89.1 % | 87.2 % |
The cations with the lowest structure matching (elements with at least 100 materials):
| site | lowest | highest |
|---|---|---|
| A | Os 82 %, Cs 84 %, Ir 85 %, Rb 85 %, B 85 %, K 87 % | Sn 95 %, Bi 94 %, Sb 94 %, Ga 94 % |
| B | Ca 82 %, Hf 85 %, Mn 86 %, B 86 %, Ta 86 %, La 87 % | Cu 94 %, Ni 94 %, Sb 93 %, Ga 93 % |
What the shared space contains¶
The modalities are aligned before the projection heads¶
The contrastive loss of this model acts on the normalised encoder outputs. There the cosine is 0.954 ± 0.011. After the projection heads, which feed the decoder, the same pairs have a cosine of -0.978 ± 0.002: the two projected vectors point in nearly opposite directions, and the decoder reads their mean.
In the benchmark
The representation task compares the two latents in the space where each model aligns them, declared per method. The cosine in both spaces is recorded on each method's page.
The 128-number latents use about two directions¶
The participation ratio, the effective number of directions a set of vectors uses, is 2.2 ± 0.0 for the structure latents and 2.3 ± 0.1 for the property latents, out of 128; the joint latent that the decoder reads uses 8.5 ± 0.8.
What this means for you
Two scalar properties can organise two directions, and the alignment pulls the structure latent onto them. A richer second modality, such as a diffraction pattern or a density of states, is the way to a richer shared space (roadmap).
The property modality of Perov-5 is thin¶
96.1 % of the materials (18,193) have a direct band gap of 0, and the 18,928 materials share only 903 distinct pairs of (formation enthalpy, band gap). 94 % of the materials have the same pair as at least ten others; the largest group has 261.
Retrieval works in one direction¶
| top 1 | median rank | |
|---|---|---|
| from a structure, find its properties | 0.230 | 186 |
| from the properties, find the structure | 0.003 | 228 |
Ranks are among all 18,928 materials; a material sharing its properties with others is not counted against. From properties to a structure, many crystals are equally right answers.
What this means for you
This is why inverse design in MEIDNet is a search and not a lookup: a material family fixes the sites, rules remove impossible compositions, and the search moves through the latent space (How MEIDNet works).
Less stable materials align better and reconstruct worse¶
By fifths of the formation enthalpy, from the lowest to the highest:
| formation enthalpy up to (eV/atom) | 0.88 | 1.22 | 1.54 | 2.00 | 5.16 |
|---|---|---|---|---|---|
| cosine | 0.952 | 0.948 | 0.950 | 0.955 | 0.963 |
| structure matching | 91.6 % | 91.6 % | 91.2 % | 89.9 % | 86.0 % |
The correlation between a material's cosine and its formation enthalpy is 0.47 ± 0.06 over the seven seeds.
What a structure alone predicts¶
On the 3,786 held-out materials, with the property of the five nearest training materials in the structure latent (three seeds):
| formation enthalpy (eV/atom) | direct band gap (eV) | |
|---|---|---|
| five nearest neighbours in the structure latent | 0.060 ± 0.004 | 0.087 ± 0.003 |
| nearest property latent (cross-modal) | 0.062 ± 0.004 | 0.131 ± 0.016 |
| linear model on the structure latent | 0.144 ± 0.015 | 0.152 ± 0.023 |
| mean of the training materials | 0.563 | 0.167 |
Mean absolute errors. The band-gap numbers are small because most gaps are zero.
What this means for you
The properties decoded from the joint latent have an error of only 0.008 eV/atom, but the joint latent contains the true properties: that is a reconstruction. The error to expect for a new structure is the first row, about 0.06 eV/atom.
In practice¶
What the results above suggest for anyone training or using such a model:
- Keep the curriculum. It is what teaches the model to rebuild crystals.
- Train three seeds, not one. Results differ between seeds, and their errors are independent.
- Choose the checkpoint on a validation measure, not at the last epoch.
- Decode with the most likely element, and add the probabilities of the seeds when you have several.
- Rank the decoded crystals by confidence (the lowest top probability over the sites) and check the composition first.
- Confirm with an MLIP, then DFT. The guide runs the first of these steps.
Steps 2, 4 and 5 are measured here on reconstruction, with the alignment model. They are not yet options of meidnet generate.