A proposed scientific-AI environment could evaluate models used for experimental prioritisation, spectral interpretation, anomaly detection, process prediction or operational decision support against a declared context of use.
Evaluation would cover provenance, leakage, calibration, uncertainty, subgroup or domain performance, distribution shift, adversarial or missing inputs, human factors and retirement. Statistical performance alone would not establish scientific validity or regulated acceptability.
