Generate highly predictive AI ADME models from your proprietary data

Cresset’s AI ADME model building technology enables the generation of reliable, highly predictive ADME models from proprietary and public data.
SHARE

Absorption, Distribution, Metabolism and Excretion (ADME) are crucial properties in the drug discovery process: even molecules with perfect activity and selectivity profile will be discarded if they show an unsatisfactory ADME profile.

Traditionally, the only way of determining ADME properties was through experiments conducted in the laboratory: however, running ADME assays for large numbers of molecules is time and resource expensive. While laboratory experiments remain the gold standard, in silico ADME models can provide a way to quickly and cost effectively identify compounds with the most promising properties that can then be experimentally investigated. This minimizes unnecessary experimental time and introduces significant cost savings, with an enhanced success rate of drug discovery. Furthermore, the addition of calculated ADME properties to Generative AI models can ensure that these properties can be directly accounted for and targeted within the generative pipelines.

Cresset’s AI ADME model building technology enables the generation of reliable, highly predictive ADME models from proprietary and public data. Predictions are associated to a Confidence Score which correlates reliably with property prediction accuracy, enabling users to make informed decisions to optimize their drug discovery processes.

Problem framing

The ADME prediction problem can be viewed in its simplest form as follows: for a given molecule, the task is to predict the value of a property of interest, which can be either a continuous value (regression) or a category (classification). In most cases, only a SMILES string and the corresponding experimental values are required to train an ADME model. Any additional molecular representations such as 3D conformers or graph-based features are generated automatically within the modelling pipeline during featurization.

The success of machine learning methods is heavily dependent on both the amount, quality, and homogeneity of training data available. Within the field of ADME predictions, this presents a considerable challenge. Publicly available data is often limited in quantity, inhomogeneous, and of varying quality. On the other hand, private datasets are of higher quality but may nevertheless be limited in size and still come with inherent errors relating to experimental measurements. In both cases, new project compounds are typically found in unexplored regions of chemical space, making it challenging to reliably predict their ADME properties.

AI ADME model building technology

Cresset’s AI ADME model building technology uses machine learning to automate the selection of the most appropriate model architecture for each ADME endpoint. Augmented with technical and scientific support from Cresset Discovery Services, this ensures that the best possible model is generated for each particular use case and data availability.

Model choice and training leverage both cutting edge and traditional approaches to ADME property prediction. As base learner models, depending on the data available and use case, we use a variety of models including gradient boosted trees, deep neural networks, and graph-based approaches. For each base learner, we use a wide range of featurization approaches: from tabular based approaches such as RDKit, to 2D and 3D graphs, to NLP based learned features. This ensures that as much chemical and physical information is provided to the models as possible.

On the side of experimental ADME data required for training the models, the distribution and type of data is considered. For classification, this includes an investigation of the balance or imbalance of the available data, and the application of data balancing techniques such as class weighting, and over and under sampling. For regression, the distribution of values is considered, with appropriate rescaling applied to ensure the best model performance possible. Censored data (labels with only an upper or lower bound rather than a point measurement) is also considered. The data is cleaned and curated to ensure units are standardized and outliers that are caused by inaccurate or mislabeled data points are removed.

On top of the selected base learners, we use Cresset’s proprietary Superlearner AI approach to combine the base learner predictions to provide a prediction with an associated Confidence Score for both regression and classification models. This approach maximizes model performance and generalizability, reducing model error and overfitting. The Confidence Score enhances the usability of the predictions by providing additional information as to how confident the model is in the prediction.

Performance evaluation

We have thoroughly benchmarked the performance of AI ADME on both public data and private data sources against various standard and cutting-edge model architecture. We have held to industry best practices to ensure that our tests are accurate with splitting strategies employed to ensure no leakage between training and test sets. An example of the performance of our framework against Chemprop is shown below.

Cresset AI ADME shows significantly higher performance vs. Chemprop on multiple ADME endpoints. Chemprop is a well-established opensource software containing message passing neural networks for molecular property prediction, widely used in pharma for ADME prediction.

ADME predictions you can trust

Cresset AI ADME is cloud native technology and can be deployed on major cloud environments. Built with scalability, parallelization, and security by design, Cresset AI ADME supports efficient training and inference (prediction) across diverse datasets and ADME endpoints. Predictions are accessible through a REST API, enabling seamless integration with existing systems and interfaces.

Our experienced Cresset Discovery Services team is on hand to provide technical and scientific support, ensuring smooth integration and helping you build ADME models that deliver strong performance for your specific use case and data availability.

Learn more about Cresset’s AI-Powered ADME model building and contact us to discuss how our technology can support your projects.

Related Science Resources

AI ADME: Build predictive models from your own data
Learn more about how to turn proprietary data into highly predictive, confidence scored models
Webinar: Revolutionizing Drug Discovery with AI
Explore how Cresset AI is revolutionizing the drug discovery process by optimizing workflows, enhancing productivity and empowering decisions...

Subscribe & Don't Miss Out

Receive our newsletter to be among the first to hear about product releases, case studies, opinion articles, events and more.