Contribute a method¶
The benchmarks rank methods under fixed protocols. Any model can take part: the protocol defines the data, the targets and the metrics, and the scoring code evaluates the outputs of a model whatever produced them. Submissions go through GitHub, so every number has a public record and a reviewer.
1. Produce the outputs¶
Each task of a protocol asks for one kind of output. For Perov-5 (protocol perov5-v1.1):
| task | output | format |
|---|---|---|
| property prediction | a prediction for every test material | predictions.csv: id (the dataset's material_id), pred_heat_all, pred_dir_gap |
| inverse design | 54 candidate structures: 6 per target and chemical family | candidates.csv: id, variant, target, formula, elements, cif; the CIFs next to it |
| representation | latent vectors of the test materials | computed by run for MEIDNet checkpoints; for other models, contact the maintainer |
The targets, the families and the budget are listed in the protocol.
2. Score them¶
python scripts/benchmarks.py score perov5 --task property_prediction --predictions predictions.csv
python scripts/benchmarks.py score perov5 --task inverse_design --candidates my_run/candidates.csv
score prints the metrics of the task as JSON. Inverse design relaxes every candidate with MACE-MP-0 first
(stability.csv and scored.csv are written next to candidates.csv); if MACE is installed in another Python
environment, pass --mlip-python or set MEIDNET_MLIP_PYTHON.
A MEIDNet checkpoint can run every task at once:
python scripts/benchmarks.py run perov5 --method meidnet-2k
3. Write the record¶
Copy benchmarks/submissions/_template.json to benchmarks/submissions/<dataset>/<method>.json and fill it in:
| field | what to put |
|---|---|
protocol |
the dataset's protocol version, for example perov5-v1.1 |
method |
name, type (model or baseline), a one-sentence description, parameters, training data, links to paper, code and weights |
method.test_in_training |
true if the training data included the test split; the row is then marked, because its test scores are not held-out |
method.alignment_space |
for the representation task: encoder or projection, the space in which the model's alignment loss compares the two latents |
results |
per task, the JSON printed by score |
spread, n_runs |
optional: the standard deviation of each metric over several training runs (seeds), and their number; results then holds the means |
artifacts |
per task, a public link to the folder with the outputs (predictions or candidates with their CIFs) |
validation_level |
mlip_validated for inverse design scored with the MLIP; dft_validated only with public DFT outputs in evidence |
status |
community_submitted |
Then run python scripts/benchmarks.py validate, which checks every task and metric name against the protocol, and
open a pull request (the template has a checklist). The maintainer re-scores the outputs with the same code; when
the numbers agree, the record is marked as computed here.
If you prefer not to open a pull request, the issue template Benchmark submission asks for the same information.
Add a dataset¶
A dataset becomes a benchmark when its file in benchmarks/datasets/ has a protocol section:
- Data and splits. Where the data live (
data_dir, withtrain.csv,val.csvandtest.csv) and which split each task uses. Results are only compared within one dataset and one protocol version. - Tasks. For each task: an
id, aname, a one-sentencesummary, itssettings(for inverse design: the family, its variants, the property targets, the candidate budget and the stability threshold) and theheadlinemetric that orders the leaderboard. - Metrics. For each metric: an
id, a shortlabel, theunit, whether higher or lower isbetter, the number ofdigitsshown and a one-sentencedefinition. Metrics marked"show": falseappear on each method's page only.
benchmarks/datasets/perov5.json is the template. Changing the targets, the budget or a definition makes a new
protocol version; earlier records keep theirs.
Common mistakes¶
- A different split. The protocol fixes the split; numbers on another split are not comparable. A model trained on all
materials can be listed, with
test_in_trainingset. - One seed. Results differ between training seeds (how much); give the mean and the spread over several runs where you can.
- A metric computed differently. Use
score, which implements the definitions of the protocol. - Outputs that are not public. A record without its outputs cannot be re-scored and is not ranked.
- A validation level the evidence does not support.
dft_validatedneeds the DFT outputs.
Submissions are reviewed by the author of MEIDNet.