Science & Validation

Built Around Scientific Reliability

MolVerity AI evaluates more than predictive accuracy. Predictive discrimination alone is not sufficient for scientific decision support; calibration, uncertainty, chemical-domain coverage and robustness under distribution shift also matter.

Current Evidence Base

Toxicity First. Broader ADMET Is Planned.

The present MolVerity research framework has been developed and benchmarked primarily for toxicity prediction. Six endpoints are currently implemented in ToxVerity. Broader ADMET functionality represents planned platform expansion and should not be interpreted as currently validated production capability.

Additional endpoints are incorporated only when suitable data, model development, validation, calibration, applicability-domain assessment and decision criteria support responsible use.
Methodological Principles

How Models Are Developed and Evaluated

01

Endpoint-specific modeling

Each implemented endpoint is modeled to reflect its biological endpoint, dataset and intended decision context.

02

Scaffold-aware train/test separation

Evaluation accounts for structural scaffolds to reduce overly optimistic estimates.

03

Repeated evaluation

Models are assessed across repeated runs rather than relying on one split.

04

Probability calibration

Raw model outputs are calibrated to improve probabilistic interpretability.

05

Uncertainty estimation

Ensemble disagreement is quantified alongside predictions.

06

Applicability-domain assessment

Each prediction is evaluated against represented chemistry.

07

Out-of-distribution detection

Molecules outside the model-development domain are flagged.

08

Explicit abstention

The platform can decline a forced prediction when support is insufficient.

09

External validation

Model behavior is assessed on held-out external data where available.

10

Bootstrap confidence intervals

Performance estimates include uncertainty ranges.

11

Transparent limitations

Known endpoint and model limitations are documented.

12

Reproducibility & versioning

Model releases and predictions are designed for traceability.

Machine Learning & Model Architectures

Multiple Model Families. One Reliability Framework.

MolVerity combines classical machine learning, neural networks and graph-based deep learning for molecular property prediction. The MolVerity framework evaluates multiple model families for each endpoint rather than assuming that a single architecture is optimal across all prediction tasks. Model selection is guided not only by predictive performance, but also by calibration, uncertainty, applicability domain and out-of-distribution behavior to support reliable molecular screening.

Three molecular machine-learning model families converging into one MolVerity reliability framework with calibration, uncertainty, applicability-domain, out-of-distribution and scaffold-context checks.

Classical QSAR Machine Learning

Molecular fingerprints are evaluated with established machine-learning methods including logistic regression, random forest, support vector machines, histogram gradient boosting and XGBoost where available. These models provide endpoint-specific QSAR benchmarks and comparative baselines.

Fingerprint Neural Networks

Multilayer perceptron neural networks learn nonlinear relationships from molecular fingerprint representations. Both single-task and multitask architectures are evaluated to study endpoint-specific performance and shared representation learning.

Graph Neural Networks

Molecular structures can also be represented directly as graphs of atoms and bonds. The MolVerity research framework includes Graph Isomorphism Networks (GIN) and graph-attention architectures based on GATv2, with both single-task and multitask implementations.

Model Selection & Reliability

Model architecture alone does not determine whether a prediction should be trusted. Candidate models are evaluated alongside probability calibration, ensemble uncertainty, applicability-domain coverage, molecular similarity, scaffold novelty, out-of-distribution detection, abstention criteria and external validation where suitable data are available.

These are model families evaluated within the MolVerity research framework. The selected calibrated model can differ by endpoint; this should not be interpreted as every deployed endpoint using every architecture listed above.
Documentation

Scientific Documentation

Formal documentation will be published as it becomes available. This page does not present fabricated performance claims.

Publications

To be published.

Validation Reports

To be published.

Model Cards

To be published.

Datasets

To be published.

Benchmark Results

To be published.

Scientific Documentation

To be published.

Scientific Collaboration

Talk to the Science Team

Reach out for methodology questions, academic collaboration or available documentation.