Using different Models
Although mother is build around catboost it basically supports other models from the ML community (like all sklearn estimators). For example, RandomForest is already supported. Furthermore, own models can be provided.
Currently supported models
Mother discovers available model wrappers from src/mother/ml/models/m_*.py.
At the moment, the built-in algorithm groups are:
catboostrandomforestlassotabpfn
You can always verify what is available in your environment:
Built-in model classes
| Algorithm key | Main model classes |
|---|---|
catboost |
CatboostRegressorMother, CatboostGaussianProcessRegressorMother, CatboostClassifierMother, CatboostRankerMother |
randomforest |
RandomForestRegressorMother, RandomForestClassifierMother |
lasso |
LassoRegressorMother, LassoClassifierBinaryMother, LassoClassifierMulticlassMother |
tabpfn |
TabPFNRegressorMother, TabPFNClassifierMother |
Easy usage patterns
The easiest way is to ask Mother for the model class by algorithm and task type, then instantiate it.
from mother import ml
# Regressor (catboost)
reg_cls = ml.get_model_class_by_algorithm_and_type("catboost", "regression")
reg = reg_cls()
# Classifier (random forest)
clf_cls = ml.get_model_class_by_algorithm_and_type("randomForest", "classification")
clf = clf_cls()
Use an explicit subtype when needed:
from mother import ml
# Lasso multiclass classifier
lasso_multi_cls = ml.get_model_class_by_algorithm_and_type(
"lasso", "classification_multiclass"
)
lasso_multi = lasso_multi_cls()
Or retrieve all classes for one algorithm and pick one:
from mother import ml
all_catboost_models = ml.get_model_class_by_algorithm("catboost")
print([m.__name__ for m in all_catboost_models])
Ranking with CatboostRankerMother
CatboostRankerMother wraps CatBoost's learning-to-rank (CatBoostRanker) for tasks where the goal is to
order items within a group (e.g. rank candidates within one experiment batch) rather than predict an
isolated value per item. Every ranking row must carry a group_id identifying which group it belongs to;
ranks are only ever compared within a group, never across groups.
import numpy as np
import pandas as pd
import sklearn
from mother.ml.models.m_catboost import CatboostRankerMother
sklearn.set_config(enable_metadata_routing=True)
rng = np.random.default_rng(0)
X_features = pd.DataFrame(rng.random((12, 3)), columns=["f0", "f1", "f2"])
y = pd.Series(rng.random(12))
groups = pd.Series([0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2])
ranker = CatboostRankerMother(logging_level="Silent", num_trees=200).set_fit_request(group_id="group_id")
ranker.fit(X=X_features, y=y, group_id=groups)
By default CatBoost uses the YetiRank loss. Other ranking losses (YetiRankPairwise, PairLogit,
PairLogitPairwise, QueryRMSE, QuerySoftMax) can be set via loss_function, or left to Optuna to choose
automatically during hyperparameter tuning. Two convenience parameters are folded into the loss_function
string automatically:
top: restrictsNDCG/MAP-style losses to the top-kitems per group (any mode exceptClassic).max_pairs: caps how many item pairs thePairLogitfamily samples per group, to keep training fast on large groups.
Predicting and estimating uncertainty per group
Because ranks and rank uncertainty only make sense within a group, use the module-level helpers instead of
calling predict/predict_uncertainty directly on a mixed-group X:
from mother.ml.models.m_catboost import (
ranker_predict_for_groups,
ranker_predict_uncertainty_for_groups,
)
ranks = ranker_predict_for_groups(ranker, X_features, group_id=groups, use_ranks=True)
uncertainty = ranker_predict_uncertainty_for_groups(ranker, X_features, group_id=groups, use_ranks=True)
predict_uncertainty estimates ranking confidence using CatBoost's virtual ensembles: several slightly
different versions of the trained model are compared, and how much they disagree on an item's score/rank
indicates how stable that item's ranking is. mother.ml.utils.groupwise_topk_analysis builds on this to flag
which top-k selections are unstable across groups.
See the ranking tutorial notebook for a full worked example, including groupwise uncertainty and out-of-fold validation.
Tip
If you are unsure about exact class names or capabilities, use:
Prediction and Uncertainty Interface
Starting with the 1.0.1 release line, model wrappers expose a more consistent prediction interface:
predict(...)returns aligned outputs across model backends.predict_uncertainty(...)provides uncertainty outputs for models that support it.
For regression models with uncertainty support, output columns follow a common naming pattern:
predictionuncertainty_datauncertainty_knowledgeuncertainty_total
Depending on model capabilities, one or more uncertainty columns can be present. The returned DataFrame keeps index alignment with the input rows to make downstream merging and analysis robust.
Note
mother_cv now has improved return typing and estimator-return behavior to better support workflows that inspect trained estimators after CV.
Tip
['LassoClassifierBinaryMother', 'LassoClassifierMulticlassMother', 'LassoRegressorMother', 'CatboostClassifierMother', 'CatboostGaussianProcessRegressorMother', 'CatboostRankerMother', 'CatboostRegressorMother', 'RandomForestClassifierMother', 'RandomForestRegressorMother']
RandomForestClassifierMother
A RandomForest classifier pipeline for the MOTHER framework, integrating hyperparameter optimization via Optuna and providing default parameter management. Inherits from both scikit-learn's RandomForestClassifier and the AbstractMotherPipeline for seamless integration with the MOTHER machine learning workflow.
get_hyperparameter_space
Defines the hyperparameter search space for RandomForestClassifier using Optuna.
Args: X: Feature matrix for training data. y: Target vector for training data. trial (Trial): Optuna trial object for suggesting hyperparameters. prefix (str, optional): Prefix to add to hyperparameter names. Defaults to "".
Returns: dict: Dictionary of hyperparameter names (with prefix) and their suggested values.
default_parameters
Returns the default hyperparameters for the RandomForestClassifier.
Args: prefix (str, optional): Prefix to add to hyperparameter names. Defaults to "".
Returns: dict: Dictionary of default hyperparameter names (with prefix) and their values.
For more information on the parent class just use 'help(ml.get_model_class("RandomForestClassifierMother")'
Using Lasso with Hyperparameter Tuning
Providing your own Model
To provide your own model and make this step as easy as possible, we provide the AbstractMotherPipelineClass.
Bases: ABC
The abstract Mother pipeline is a conventional sklearn estimator / transformer etc. but adds methods for hyperparameter definition. Furthermore, it ensures for non sklearn classes and derived classes that they are compatible to the sklearn pipeline interface. This is done by implementing the get_params and set_params methods.
predict_uncertainty(X, **kwargs)
Coordinating method for uncertainty prediction. This is a simple fallback for models without a specialized predict_uncertainty implementation.
Models with specialized uncertainty estimation (e.g., CatBoost, TabPFN, RandomForest) should override this method with their own implementations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Input features to predict target values |
required |
**kwargs
|
Additional keyword arguments |
{}
|
Your own model just has to inherit from that class and implement the required functions that provide the hyperparameters you want to tune. For example, see the implementation of the Lasso model. Since lasso basically has one parameter to be tuned, the implementation is fairly easy.
import logging
from typing import Mapping, Optional, Union
from optuna.trial import Trial
from sklearn.linear_model import Lasso, LogisticRegression
from mother.ml.core import AbstractMotherPipeline
from mother.ml.models import utils
module_logger: logging.Logger = logging.getLogger(__name__)
class LassoRegressorMother(Lasso, AbstractMotherPipeline):
"""
MOTHER class for a LASSO regression including hyperparameter optimization
"""
def get_hyperparameter_space(self, X, y, trial: Trial, prefix: str = "") -> dict:
"""
Define the hyperparameter search space for Lasso regression.
Parameters:
X: array-like
Feature matrix.
y: array-like
Target vector.
trial: optuna.trial.Trial
Optuna trial object for suggesting hyperparameters.
prefix: str, optional
Prefix to add to hyperparameter names.
Returns:
dict: Dictionary containing hyperparameter names and their suggested values.
"""
return utils.add_prefix_to_dict_keys(
{"alpha": trial.suggest_float(prefix + "alpha", 1e-6, 1e1, log=True)},
prefix=prefix,
)
def default_parameters(self, prefix: str = "") -> dict:
"""
Return the default hyperparameters for the Lasso model.
Parameters:
prefix: str, optional
Prefix to add to hyperparameter names.
Returns:
dict: Dictionary containing default hyperparameter values.
"""
return utils.add_prefix_to_dict_keys({"alpha": 1e-3}, prefix=prefix)
def set_params(self, **params):
"""
Set the parameters of the Lasso model.
Parameters:
**params: Keyword arguments for the parameters to set.
"""
return super().set_params(**params)
def get_params(self, deep=True) -> dict:
return super().get_params(deep=deep)
class LassoClassifierBinaryMother(LogisticRegression, AbstractMotherPipeline):
"""
MOTHER class for a LASSO classification including hyperparameter optimization
"""
def __init__(
self,
*,
dual: bool = False,
tol: float = 0.0001,
C: float = 1,
fit_intercept: bool = True,
intercept_scaling: float = 1,
class_weight: Optional[Union[Mapping, str]] = "balanced",
random_state: int = 42,
solver: str = "liblinear", # liblinear supports pure L1 (l1_ratio=1)
max_iter: int = 3000,
verbose: int = 0,
warm_start: bool = False,
n_jobs: Optional[int] = None,
) -> None:
# l1_ratio=1: pure L1 regularisation via the sklearn ≥1.8 API.
# penalty='l1' was deprecated in sklearn 1.8 and will be removed in 1.10;
# l1_ratio=1 is the replacement (analogous to LogisticRegressionCV).
super().__init__(
l1_ratio=1,
dual=dual,
tol=tol,
C=C,
fit_intercept=fit_intercept,
intercept_scaling=intercept_scaling,
class_weight=class_weight,
random_state=random_state,
solver=solver, # type: ignore
max_iter=max_iter,
verbose=verbose,
warm_start=warm_start,
n_jobs=n_jobs,
)
def get_hyperparameter_space(self, X, y, trial: Trial, prefix: str = "") -> dict:
"""
Define the hyperparameter search space for Lasso classification.
Parameters:
X: array-like
Feature matrix.
y: array-like
Target vector.
trial: optuna.trial.Trial
Optuna trial object for suggesting hyperparameters.
prefix: str, optional
Prefix to add to hyperparameter names.
Returns:
dict: Dictionary containing hyperparameter names and their suggested values.
"""
return utils.add_prefix_to_dict_keys(
{
"C": trial.suggest_float(prefix + "C", 1e-6, 1e1, log=True),
},
prefix=prefix,
)
def default_parameters(self, prefix: str = "") -> dict:
"""
Return the default hyperparameters for the Lasso model.
Parameters:
prefix: str, optional
Prefix to add to hyperparameter names.
Returns:
dict: Dictionary containing default hyperparameter values.
"""
return utils.add_prefix_to_dict_keys({"C": 1e0}, prefix=prefix)
def set_params(self, **params):
"""
Set the parameters of the Lasso model.
Parameters:
**params: Keyword arguments for the parameters to set.
"""
return super().set_params(**params)
def get_params(self, deep=True) -> dict:
return super().get_params(deep=deep)
class LassoClassifierMulticlassMother(LassoClassifierBinaryMother):
"""
MOTHER class for a LASSO classification with multiclass support.
Inherits from LassoClassifierBinaryMother.
"""
def __init__(self, **kwargs):
module_logger.warning(
"""LassoClassifierMother selected. 'Saga' is used as solver.
Scale input features beforehand to improve convergence."""
)
if "solver" in kwargs:
module_logger.warning(
"LassoClassifierMulticlassMother selected. 'Saga' is used as solver for multiclass problems."
)
# Use 'saga' solver for multiclass support
# 'saga' supports L1 penalty and is suitable for large datasets
kwargs["solver"] = "saga"
super().__init__(**kwargs)
Registering your model using MotherModelRegistry
To register your own model and use it easily within the mother framework you can register your model with the available decorator.
MotherModelRegistry
Singleton registry for dynamically discovering and managing model classes in the 'mother.ml.models' package.
This class scans the models directory for Python files matching the pattern 'm_*.py', imports them, and registers all classes that inherit from AbstractMotherPipeline. It provides mappings for model class names, lower-case lookups, and a list of supported algorithms. The registry is used to facilitate model discovery, retrieval, and algorithm support checks throughout the mother.ml package.
Attributes:
| Name | Type | Description |
|---|---|---|
models_dir |
Path
|
Path to the directory containing model modules. |
model_classes |
dict
|
Mapping of model class names to their class objects. |
model_classes_lower |
dict
|
Mapping of lower-case model class names to their canonical names. |
supported_algorithms |
list
|
List of supported algorithm names discovered from model files. |
list_registered_models()
List all registered models with their algorithms.
register_model(model_class, algorithm=None)
Register a model class manually.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_class
|
Type[AbstractMotherPipeline]
|
The model class to register |
required |
algorithm
|
Optional[str]
|
Optional algorithm name. If not provided, will be derived from class name |
None
|
unregister_model(model_class_name)
Unregister a model class.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_class_name
|
str
|
Name of the model class to unregister |
required |
algo_is_supported(algorithm)
describe_model(name)
get_available_algorithms()
get_model_class(name)
Get a model class by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Name of the model class |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Type |
type[AbstractMotherPipeline]
|
The model class |
Raises:
| Type | Description |
|---|---|
KeyError
|
If the model class is not found |
get_model_class_by_algorithm(algorithm)
Get model classes by algorithm name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
algorithm
|
str
|
Name of the algorithm |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Type |
list[Type[AbstractMotherPipeline]]
|
The model class |
get_model_class_by_algorithm_and_type(algorithm, model_type)
Returns the appropriate model class based on the algorithm and model type.
get_supported_models()
See the following example how to implement your custom RandomForest Classifier.
Example
from mother import ml
from sklearn.ensemble import RandomForestClassifier
@ml.register_model("custom_rf")
class CustomRandomForestMother(RandomForestClassifier, ml.AbstractMotherPipeline):
def get_hyperparameter_space(self, X, y, trial, prefix=""):
return {
f"{prefix}n_estimators": trial.suggest_int("n_estimators", 10, 100),
f"{prefix}max_depth": trial.suggest_int("max_depth", 3, 10),
}
def default_parameters(self, prefix=""):
return {f"{prefix}n_estimators": 50, f"{prefix}max_depth": 5}
print(ml.get_model_class_by_algorithm("custom_rf"))