MEIDNet Benchmarks¶
Standardised tasks for multimodal materials models, each with a fixed protocol: the data split, the targets, the candidate budget and the metrics. Every method is evaluated the same way, the outputs behind each number are kept, and the scoring works for any model, not only MEIDNet. Results are compared within one dataset and one protocol version.
1protocols
3tasks
10methods
11computed here
Leaderboards¶
| dataset | protocol | tasks | models | baselines |
|---|---|---|---|---|
| Perov-5 | perov5-v1.1 |
inverse design, property prediction, representation | 5 | 5 |
| Carbon-24 | not yet defined | – | – | – |
| MP-20 | not yet defined | – | – | – |
Analysis¶
What the benchmark runs show beyond one number per method: the spread between training seeds, results on materials held out of training, where reconstruction fails, and what the shared space contains.
How a result gets on a leaderboard¶
- Protocol. Each dataset fixes its tasks in
benchmarks/datasets/<dataset>.json: splits, property targets, candidate budget, metrics and the direction in which each is better. - Run or score.
python scripts/benchmarks.py runevaluates a MEIDNet checkpoint end to end;python scripts/benchmarks.py scoreevaluates the predictions or candidate structures of any other model with the same code. - Record. One JSON file per method (
benchmarks/submissions/<dataset>/<method>.json) holds the description, the numbers per task and the folder with the outputs behind them. - Verification. Numbers computed in this repository carry a record in
benchmarks/verified/; submitted results are re-scored from their outputs before they are marked as computed here.