Contribute a benchmark result¶
MEIDNet Benchmarks collect results per dataset, each with the evidence behind it. Anyone who runs MEIDNet on a dataset — the published Perov-5, or their own — can submit a result. Submissions go through GitHub, so every number has a public record and a reviewer.
What a submission is¶
One JSON file, benchmarks/submissions/<dataset>/<id>.json, that follows benchmarks/schema.json. The short
version:
| field | what to put |
|---|---|
dataset |
an id from benchmarks/datasets/ (perov5, …); a new dataset needs its own datasets/<id>.json with source, size and split |
modalities |
what the model aligned: structure, property:<column> (one per property), later xrd, … |
model |
MEIDNet version, family and variant, the configuration (meidnet.yaml), checkpoint link and its SHA-256 |
category |
representation_quality, conditional_design, efficiency or validation — one category per file |
results |
metric → {value, unit}; only the metric names allowed for that category (listed in the schema) |
validation_level |
the evidence ladder: generated → ml_filtered → mlip_validated → dft_validated → experimentally_validated |
evidence |
public links: CIF archive, stability.csv, DFT inputs/outputs, paper DOI |
reproduce |
configuration, seed and the exact command that produces the numbers |
status |
always community_submitted for a new file |
The five things people get wrong¶
- Mixing categories in one file. A representation result (MAE, R², retrieval) and a design result (SUN rate) are two files.
- A metric name that is not in the list.
validaterefuses it; the lists exist so that rows are comparable. - No split. State the split you evaluated on (or link a file of ids). A number on an unknown split is not comparable.
- Evidence that is not public. A result without an inspectable artefact stays at
generated. - Claiming a validation level the evidence does not support.
dft_validatedneeds the DFT outputs, not the intention.
How it works¶
- Copy
benchmarks/submissions/_template.json, fill it in, runpython scripts/benchmarks.py validate. - Open a pull request (the template has a checklist). CI runs
validate. - Once merged, the result is rendered on the dataset's benchmark page and gets a permanent page of its own, marked Community submitted.
- The maintainer may run
python scripts/benchmarks.py reproduce <id>: it re-runs the stated command from the stated configuration and compares the numbers. If they agree within the tolerances, the row becomes MEIDNet verified (a record inbenchmarks/verified/, which only the maintainer can add).
Verified means reproduced from the submitted configuration with this code. It does not mean the material is
real: that is what the validation level says. A re-run that does not agree is kept too
(benchmarks/verified/<id>.attempt.json) and shown on the row's page next to the submitted numbers, so that a
disagreement is visible rather than silent — this applies to the paper's own rows as much as to anyone else's.
Representation rows that are reproduced here also get their per-material predictions saved
(benchmarks/verified/<id>.predictions.csv), which the pages turn into parity plots and a latent-agreement histogram.
If you prefer not to open a pull request, the issue template Benchmark submission asks for the same fields and the maintainer converts it. Submissions are reviewed by the author of MEIDNet; there is no promised response time.