Preliminary local QM9 point estimates · candidate split · redox explorer is synthetic

Open benchmark · molecular property transfer

Transfer—or expensive similarity?

Read the first aggregate QM9 comparison, then explore why molecular train/test splits change what a model appears to know. The longer-term question is whether pretrained representations transfer beyond nearby chemical structures.

This static page does not run inference. The QM9 section reads a tracked, aggregate-only result from one completed local run; its split and uncertainty caveats remain open. The redox explorer uses tiny, invented fixture values only.

QM9-28M · preliminary local result

Lower error on this candidate split.

AGGREGATE ONLY no rows, labels, predictions, or weights

The released, already fine-tuned MIST-28M predictor had lower error than the locked ECFP4 Ridge baseline on all 12 QM9 targets. This comparison uses DFT-computed labels and a split reconstructed from public code; it is not an official checkpoint reproduction or a causal test of pretraining.

Released MIST mean normalized MAE
Locked ECFP Ridge mean normalized MAE
Error reduction relative to Ridge
Selected cohort test rows
HOMO, LUMO, and gap MAE Shorter is better · one shared hartree scale
All 12 targets for the selected cohort
Target Unit MIST MAE Ridge MAE MAE reduction MIST R² Ridge R²

Mean normalized MAE averages each target's MAE divided by its training-set standard deviation. Positive reduction values favor MIST. Values are descriptive point estimates; the tracked JSON retains full precision.

Loading authenticated aggregate provenance…

Separate synthetic redox track

Change the test, change the story.

DEMO synthetic software output

Synthetic software-demo outputs

Loading fixture…

MAEvolts · test
RMSEvolts · test
descriptive only
Median similarityECFP Tanimoto
Predicted vs. synthetic target Test-set rows only
synthetic test row ideal prediction
Test-row details for the selected synthetic run
Fixture ID Family Target (V) Prediction (V) Abs. error (V) Nearest similarity

Do not rank models from these numbers. The tiny targets are invented, sample sizes are intentionally inadequate, and MIST is absent. This view verifies the interface, split logic, provenance, and reporting path only.

How to read the benchmark

A result is only as honest as its holdout.

A random split asks whether a model can interpolate. Harder splits ask whether useful structure survives when familiar scaffolds, families, or sources disappear.

We never compare a raw foundation checkpoint with a trained model. The planned MIST arm starts from an existing pretrained checkpoint, then attaches and trains a regression head on the same labeled redox rows as the ECFP models. Optional LoRA updates are trained alongside that head. Both final predictors are compared on the same held-out measurements.

01

Freeze the question

Use one target definition, one condition cohort, and immutable molecule identities.

02

Stress the split

Move whole molecular groups, then audit nearest-training similarity for every test row.

03

Earn the claim

Compare the downstream-trained MIST predictor with label-trained classical baselines.

Current scientific status

One result, not the final answer.

The preliminary QM9 predictor comparison is complete. The broader transfer claim gate remains closed: real redox data, downstream MIST training, harder holdouts, learning curves, repeated seeds, uncertainty, and an independently curated external set are still required.