Feature Generation in Cheminformatics
Feature generation is the process of extracting meaningful numerical values from chemical structures. These features capture important chemical properties and patterns, enabling machine learning models to learn relationships between molecular structure and target properties.
Chemical Features supported by Mother
The Mother framework provides tools to automate feature generation:
-
Descriptors
- Numerical values summarizing molecular properties (e.g., Molecular weight, LogP (hydrophobicity), number of rotatable bonds, topological polar surface area).
- Mother supports all descriptors provided by rdkit.Chem.Descriptors. You can check the full list of available descriptors:
-
Fingerprints
- Binary or count vectors encoding the presence or absence of specific substructures
- Mother supports
- MACCS keys
MaccsFingerprints - A convenience class
FingerprintsGenericto generate fingerprints supported byrdkit.Chem.rdFingerprintGenerator. You can provide one from the following list to thefp_typeparameter:- RDKitFP
- MorganFP
- AtomPairFP
- TopologicalTorsionFP
- MACCS keys
Binary vs. Count Fingerprints
By default,
FingerprintsGenericandMorganFingerprintsgenerate count-based vectors, which record how many times each substructure occurs in a molecule. Count fingerprints can improve model performance when substructure frequency is informative. Setuse_counts=Falseto generate a binary vector, where each bit is either0(substructure absent) or1(substructure present).Mode Parameter Output values Use case Count (default) use_counts=True0, 1, 2, ... When substructure frequency matters Binary use_counts=False0 or 1 When only substructure presence matters The
use_countsparameter is available on bothFingerprintsGenericand theMorganFingerprintsconvenience class:
from rdkit import Chem
from mother.feature_generation import FingerprintsGeneric, MorganFingerprints
molecule_objects = [Chem.MolFromSmiles(smi) for smi in ["CCO", "CCN", "c1ccccc1"]]
# Count-based Morgan fingerprints are the default
morgan_counts = MorganFingerprints(radius=2, fpSize=2048)
features = morgan_counts.fit_transform(molecule_objects)
# Count-based fingerprints via the generic class (works with any supported fp_type)
fp_counts = FingerprintsGeneric(
fp_type="AtomPairFP",
parameters={"fpSize": 2048},
)
features = fp_counts.fit_transform(molecule_objects)
# Opt in to binary Morgan fingerprints
morgan_binary = MorganFingerprints(radius=2, fpSize=2048, use_counts=False)
!!! note
`use_counts` is a standard scikit-learn parameter — it is preserved through `set_params()`, `get_params()`, and `sklearn.base.clone()`.
- Custom Features from Machine Learning Models
- Features generated by trained machine learning models
- Any input data can be provided as input to the final machine learning object.
Usage Example
from mother.feature_generation import ChemicalDescriptors
# Example: Generate molecular weight and LogP descriptors
descriptor_generator = ChemicalDescriptors(descriptor_list=["RingCount", "MolLogP"])
features = descriptor_generator.fit_transform(molecule_objects)
You can combine multiple feature generators using scikit-learn’s FeatureUnion:
from sklearn.pipeline import FeatureUnion
from rdkit import Chem
from mother.feature_generation import ChemicalDescriptors, MorganFingerprints
molecule_objects = [Chem.MolFromSmiles(smi) for smi in ["CCO", "CCN", "c1ccccc1"]]
feature_generator = FeatureUnion(
[
("descriptors", ChemicalDescriptors(descriptor_list=["MolWt", "MolLogP"])),
("morgan_fp", MorganFingerprints(radius=2, fpSize=2048)),
]
)
features = feature_generator.fit_transform(molecule_objects)