Machine learning infers molecular assembly from mass spectra as a universal biosignature
A new paper from the Cronin Group, published in PNAS, shows that molecular assembly (MA) can be inferred directly from mass spectrometry data without structural elucidation, making it a practical universal biosignature. The team trained an XGBoost model on single-stage electron ionisation spectra from the NIST Chemistry WebBook, the type of data expected from instruments proposed for missions to Titan and Europa, and predicted MA with a threefold lower error than the best baseline method. The model generalised to independent MassBank spectra and tended to underpredict high MA values, making it conservative with respect to false positives. Simulated data at different collision energies showed that small variations in instrument settings can double the prediction error, underlining the need for calibrated, well-documented spectral databases.
The paper is open access and can be read on the PNAS website.