Skip to content

MEIDNet Benchmarks

Standardised tasks for multimodal materials models, each with a fixed protocol: the data split, the targets, the candidate budget and the metrics. Every method is evaluated the same way, the outputs behind each number are kept, and the scoring works for any model, not only MEIDNet. Results are compared within one dataset and one protocol version.

1protocols
3tasks
10methods
11computed here

Leaderboards

dataset protocol tasks models baselines
Perov-5 perov5-v1.1 inverse design, property prediction, representation 5 5
Carbon-24 not yet defined – – –
MP-20 not yet defined – – –

Analysis

What the benchmark runs show beyond one number per method: the spread between training seeds, results on materials held out of training, where reconstruction fails, and what the shared space contains.

How a result gets on a leaderboard

  1. Protocol. Each dataset fixes its tasks in benchmarks/datasets/<dataset>.json: splits, property targets, candidate budget, metrics and the direction in which each is better.
  2. Run or score. python scripts/benchmarks.py run evaluates a MEIDNet checkpoint end to end; python scripts/benchmarks.py score evaluates the predictions or candidate structures of any other model with the same code.
  3. Record. One JSON file per method (benchmarks/submissions/<dataset>/<method>.json) holds the description, the numbers per task and the folder with the outputs behind them.
  4. Verification. Numbers computed in this repository carry a record in benchmarks/verified/; submitted results are re-scored from their outputs before they are marked as computed here.

Contribute a method · Databases by application