A retrospective validation of deep learning imputation for prioritising agrochemical screening

The emergence of resistance and increased stringency of regulatory requirements have created a need for new agrochemicals. The long product development process and increased costs mean there is a need to explore methods to improve the success of agrochemicals throughout the development cycle. Machine learning methods are increasingly being applied to optimise their development. Cerella™ uses a deep learning method that imputes missing data in an experimental data matrix. It accepts both molecular descriptors and sparse experimental data as input, exploiting the relationships between experimentally measured endpoints [1, 2].

An ensemble of networks generates a probability distribution for each individual prediction, accounting for uncertainties in both the experimental data and in extrapolations from the training data. From this, confidence in each prediction can be assessed.

Here, we explore how Cerella can be combined with multi-parameter optimisation (MPO) to accurately prioritise compounds with an objective of controlling a broadleaf weed species to protect corn and soybean crops.

Download poster pdf