Scope and roadmap¶
What works today¶
| Structures | any CIF with up to max_sites atoms (default 20); any elements |
| Properties | any number of scalar columns |
| Training | your data, your properties; the published settings by default |
| Generation | prototype families: a fixed arrangement of sites whose occupants and cell size are chosen. Built in: cubic ABX₃, A₂BB′X₆; write your own in YAML |
| Rules | charge balance, tolerance and octahedral factors, bond windows, minimum distance, predicted-property windows, element filters, and your own Python rules |
| Explanations | reports for data, training and generation; the Studio for live what-if |
What does not work (yet)¶
- New atomic arrangements. The decoder's positions are not used to create geometry; candidates are compositions placed on the family prototype. Generating free geometry needs a decoder trained to produce it (the published decoder receives the true positions during training).
- Spectra, images, text as modalities. The property modality is a vector of scalars. Binned XRD or DOS could be added as a vector modality with its own encoder; it is designed for but not implemented.
- Large cells.
max_sitescan be raised, but memory and time grow with the square of it, and the published model was trained on 5-atom cells. - Fine-tuning the published model on new properties. Its property head has two outputs; train a new model.
Roadmap (in order of likely value)¶
- Vector modalities (XRD, DOS) with a shared-latent encoder each.
- A
training.init_fromoption to continue training from a checkpoint with the same properties. - A geometry-generating decoder variant (trained from prototype positions instead of true positions —
model.decoder_coordinate_input: prototypeis the first step and is available as an experimental setting). - More built-in families (spinel, Heusler, rocksalt/zincblende, MXene M₂XT₂).
Contributions of family files and rules are the easiest way to help; see the GitHub repository.