Skip to content

Using different Models

Although mother is build around catboost it basically supports other models from the ML community (like all sklearn estimators). For example, RandomForest is already supported. Furthermore, own models can be provided.

Currently supported models

Mother discovers available model wrappers from src/mother/ml/models/m_*.py. At the moment, the built-in algorithm groups are:

  • catboost
  • randomforest
  • lasso
  • tabpfn

You can always verify what is available in your environment:

Python
from mother import ml

print(ml.get_available_algorithms())
print(ml.get_supported_models())

Built-in model classes

Algorithm key Main model classes
catboost CatboostRegressorMother, CatboostGaussianProcessRegressorMother, CatboostClassifierMother, CatboostRankerMother
randomforest RandomForestRegressorMother, RandomForestClassifierMother
lasso LassoRegressorMother, LassoClassifierBinaryMother, LassoClassifierMulticlassMother
tabpfn TabPFNRegressorMother, TabPFNClassifierMother

Easy usage patterns

The easiest way is to ask Mother for the model class by algorithm and task type, then instantiate it.

Python
from mother import ml

# Regressor (catboost)
reg_cls = ml.get_model_class_by_algorithm_and_type("catboost", "regression")
reg = reg_cls()

# Classifier (random forest)
clf_cls = ml.get_model_class_by_algorithm_and_type("randomForest", "classification")
clf = clf_cls()

Use an explicit subtype when needed:

Python
from mother import ml

# Lasso multiclass classifier
lasso_multi_cls = ml.get_model_class_by_algorithm_and_type(
    "lasso", "classification_multiclass"
)
lasso_multi = lasso_multi_cls()

Or retrieve all classes for one algorithm and pick one:

Python
from mother import ml

all_catboost_models = ml.get_model_class_by_algorithm("catboost")
print([m.__name__ for m in all_catboost_models])

Ranking with CatboostRankerMother

CatboostRankerMother wraps CatBoost's learning-to-rank (CatBoostRanker) for tasks where the goal is to order items within a group (e.g. rank candidates within one experiment batch) rather than predict an isolated value per item. Every ranking row must carry a group_id identifying which group it belongs to; ranks are only ever compared within a group, never across groups.

Python
import numpy as np
import pandas as pd
import sklearn

from mother.ml.models.m_catboost import CatboostRankerMother

sklearn.set_config(enable_metadata_routing=True)

rng = np.random.default_rng(0)
X_features = pd.DataFrame(rng.random((12, 3)), columns=["f0", "f1", "f2"])
y = pd.Series(rng.random(12))
groups = pd.Series([0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2])

ranker = CatboostRankerMother(logging_level="Silent", num_trees=200).set_fit_request(group_id="group_id")
ranker.fit(X=X_features, y=y, group_id=groups)

By default CatBoost uses the YetiRank loss. Other ranking losses (YetiRankPairwise, PairLogit, PairLogitPairwise, QueryRMSE, QuerySoftMax) can be set via loss_function, or left to Optuna to choose automatically during hyperparameter tuning. Two convenience parameters are folded into the loss_function string automatically:

  • top: restricts NDCG/MAP-style losses to the top-k items per group (any mode except Classic).
  • max_pairs: caps how many item pairs the PairLogit family samples per group, to keep training fast on large groups.

Predicting and estimating uncertainty per group

Because ranks and rank uncertainty only make sense within a group, use the module-level helpers instead of calling predict/predict_uncertainty directly on a mixed-group X:

Python
from mother.ml.models.m_catboost import (
    ranker_predict_for_groups,
    ranker_predict_uncertainty_for_groups,
)

ranks = ranker_predict_for_groups(ranker, X_features, group_id=groups, use_ranks=True)
uncertainty = ranker_predict_uncertainty_for_groups(ranker, X_features, group_id=groups, use_ranks=True)

predict_uncertainty estimates ranking confidence using CatBoost's virtual ensembles: several slightly different versions of the trained model are compared, and how much they disagree on an item's score/rank indicates how stable that item's ranking is. mother.ml.utils.groupwise_topk_analysis builds on this to flag which top-k selections are unstable across groups.

See the ranking tutorial notebook for a full worked example, including groupwise uncertainty and out-of-fold validation.

Tip

If you are unsure about exact class names or capabilities, use:

Python
from mother import ml
print(ml.describe_model("RandomForestClassifierMother"))

Prediction and Uncertainty Interface

Starting with the 1.0.1 release line, model wrappers expose a more consistent prediction interface:

  • predict(...) returns aligned outputs across model backends.
  • predict_uncertainty(...) provides uncertainty outputs for models that support it.

For regression models with uncertainty support, output columns follow a common naming pattern:

  • prediction
  • uncertainty_data
  • uncertainty_knowledge
  • uncertainty_total

Depending on model capabilities, one or more uncertainty columns can be present. The returned DataFrame keeps index alignment with the input rows to make downstream merging and analysis robust.

Note

mother_cv now has improved return typing and estimator-return behavior to better support workflows that inspect trained estimators after CV.

Tip

Python
from mother import ml
print(ml.get_available_algorithms())

['lasso', 'catboost', 'randomforest']

Python
from mother import ml
print(ml.get_supported_models())

['LassoClassifierBinaryMother', 'LassoClassifierMulticlassMother', 'LassoRegressorMother', 'CatboostClassifierMother', 'CatboostGaussianProcessRegressorMother', 'CatboostRankerMother', 'CatboostRegressorMother', 'RandomForestClassifierMother', 'RandomForestRegressorMother']

Python
from mother import ml
print(ml.describe_model("RandomForestClassifierMother"))

RandomForestClassifierMother

A RandomForest classifier pipeline for the MOTHER framework, integrating hyperparameter optimization via Optuna and providing default parameter management. Inherits from both scikit-learn's RandomForestClassifier and the AbstractMotherPipeline for seamless integration with the MOTHER machine learning workflow.

get_hyperparameter_space

Defines the hyperparameter search space for RandomForestClassifier using Optuna.

Args: X: Feature matrix for training data. y: Target vector for training data. trial (Trial): Optuna trial object for suggesting hyperparameters. prefix (str, optional): Prefix to add to hyperparameter names. Defaults to "".

Returns: dict: Dictionary of hyperparameter names (with prefix) and their suggested values.

default_parameters

Returns the default hyperparameters for the RandomForestClassifier.

Args: prefix (str, optional): Prefix to add to hyperparameter names. Defaults to "".

Returns: dict: Dictionary of default hyperparameter names (with prefix) and their values.

For more information on the parent class just use 'help(ml.get_model_class("RandomForestClassifierMother")'

Using Lasso with Hyperparameter Tuning

Providing your own Model

To provide your own model and make this step as easy as possible, we provide the AbstractMotherPipelineClass.

Bases: ABC

The abstract Mother pipeline is a conventional sklearn estimator / transformer etc. but adds methods for hyperparameter definition. Furthermore, it ensures for non sklearn classes and derived classes that they are compatible to the sklearn pipeline interface. This is done by implementing the get_params and set_params methods.

predict_uncertainty(X, **kwargs)

Coordinating method for uncertainty prediction. This is a simple fallback for models without a specialized predict_uncertainty implementation.

Models with specialized uncertainty estimation (e.g., CatBoost, TabPFN, RandomForest) should override this method with their own implementations.

Parameters:

Name Type Description Default
X DataFrame

Input features to predict target values

required
**kwargs

Additional keyword arguments

{}

Your own model just has to inherit from that class and implement the required functions that provide the hyperparameters you want to tune. For example, see the implementation of the Lasso model. Since lasso basically has one parameter to be tuned, the implementation is fairly easy.

Lasso with Hyperparameter Tuning
import logging
from typing import Mapping, Optional, Union

from optuna.trial import Trial
from sklearn.linear_model import Lasso, LogisticRegression

from mother.ml.core import AbstractMotherPipeline
from mother.ml.models import utils

module_logger: logging.Logger = logging.getLogger(__name__)


class LassoRegressorMother(Lasso, AbstractMotherPipeline):
    """
    MOTHER class for a LASSO regression including hyperparameter optimization
    """

    def get_hyperparameter_space(self, X, y, trial: Trial, prefix: str = "") -> dict:
        """
        Define the hyperparameter search space for Lasso regression.

        Parameters:
            X: array-like
                Feature matrix.
            y: array-like
                Target vector.
            trial: optuna.trial.Trial
                Optuna trial object for suggesting hyperparameters.
            prefix: str, optional
                Prefix to add to hyperparameter names.

        Returns:
            dict: Dictionary containing hyperparameter names and their suggested values.
        """
        return utils.add_prefix_to_dict_keys(
            {"alpha": trial.suggest_float(prefix + "alpha", 1e-6, 1e1, log=True)},
            prefix=prefix,
        )

    def default_parameters(self, prefix: str = "") -> dict:
        """
        Return the default hyperparameters for the Lasso model.

        Parameters:
            prefix: str, optional
                Prefix to add to hyperparameter names.

        Returns:
            dict: Dictionary containing default hyperparameter values.
        """
        return utils.add_prefix_to_dict_keys({"alpha": 1e-3}, prefix=prefix)

    def set_params(self, **params):
        """
        Set the parameters of the Lasso model.

        Parameters:
        **params: Keyword arguments for the parameters to set.
        """
        return super().set_params(**params)

    def get_params(self, deep=True) -> dict:
        return super().get_params(deep=deep)


class LassoClassifierBinaryMother(LogisticRegression, AbstractMotherPipeline):
    """
    MOTHER class for a LASSO classification including hyperparameter optimization
    """

    def __init__(
        self,
        *,
        dual: bool = False,
        tol: float = 0.0001,
        C: float = 1,
        fit_intercept: bool = True,
        intercept_scaling: float = 1,
        class_weight: Optional[Union[Mapping, str]] = "balanced",
        random_state: int = 42,
        solver: str = "liblinear",  # liblinear supports pure L1 (l1_ratio=1)
        max_iter: int = 3000,
        verbose: int = 0,
        warm_start: bool = False,
        n_jobs: Optional[int] = None,
    ) -> None:
        # l1_ratio=1: pure L1 regularisation via the sklearn ≥1.8 API.
        # penalty='l1' was deprecated in sklearn 1.8 and will be removed in 1.10;
        # l1_ratio=1 is the replacement (analogous to LogisticRegressionCV).
        super().__init__(
            l1_ratio=1,
            dual=dual,
            tol=tol,
            C=C,
            fit_intercept=fit_intercept,
            intercept_scaling=intercept_scaling,
            class_weight=class_weight,
            random_state=random_state,
            solver=solver,  # type: ignore
            max_iter=max_iter,
            verbose=verbose,
            warm_start=warm_start,
            n_jobs=n_jobs,
        )

    def get_hyperparameter_space(self, X, y, trial: Trial, prefix: str = "") -> dict:
        """
        Define the hyperparameter search space for Lasso classification.

        Parameters:
            X: array-like
                Feature matrix.
            y: array-like
                Target vector.
            trial: optuna.trial.Trial
                Optuna trial object for suggesting hyperparameters.
            prefix: str, optional
                Prefix to add to hyperparameter names.

        Returns:
            dict: Dictionary containing hyperparameter names and their suggested values.
        """
        return utils.add_prefix_to_dict_keys(
            {
                "C": trial.suggest_float(prefix + "C", 1e-6, 1e1, log=True),
            },
            prefix=prefix,
        )

    def default_parameters(self, prefix: str = "") -> dict:
        """
        Return the default hyperparameters for the Lasso model.

        Parameters:
            prefix: str, optional
                Prefix to add to hyperparameter names.

        Returns:
            dict: Dictionary containing default hyperparameter values.
        """
        return utils.add_prefix_to_dict_keys({"C": 1e0}, prefix=prefix)

    def set_params(self, **params):
        """
        Set the parameters of the Lasso model.

        Parameters:
        **params: Keyword arguments for the parameters to set.
        """
        return super().set_params(**params)

    def get_params(self, deep=True) -> dict:
        return super().get_params(deep=deep)


class LassoClassifierMulticlassMother(LassoClassifierBinaryMother):
    """
    MOTHER class for a LASSO classification with multiclass support.
    Inherits from LassoClassifierBinaryMother.
    """

    def __init__(self, **kwargs):
        module_logger.warning(
            """LassoClassifierMother selected. 'Saga' is used as solver.
            Scale input features beforehand to improve convergence."""
        )
        if "solver" in kwargs:
            module_logger.warning(
                "LassoClassifierMulticlassMother selected. 'Saga' is used as solver for multiclass problems."
            )
            # Use 'saga' solver for multiclass support
            # 'saga' supports L1 penalty and is suitable for large datasets
        kwargs["solver"] = "saga"
        super().__init__(**kwargs)

Registering your model using MotherModelRegistry

To register your own model and use it easily within the mother framework you can register your model with the available decorator.

MotherModelRegistry

Singleton registry for dynamically discovering and managing model classes in the 'mother.ml.models' package.

This class scans the models directory for Python files matching the pattern 'm_*.py', imports them, and registers all classes that inherit from AbstractMotherPipeline. It provides mappings for model class names, lower-case lookups, and a list of supported algorithms. The registry is used to facilitate model discovery, retrieval, and algorithm support checks throughout the mother.ml package.

Attributes:

Name Type Description
models_dir Path

Path to the directory containing model modules.

model_classes dict

Mapping of model class names to their class objects.

model_classes_lower dict

Mapping of lower-case model class names to their canonical names.

supported_algorithms list

List of supported algorithm names discovered from model files.

list_registered_models()

List all registered models with their algorithms.

register_model(model_class, algorithm=None)

Register a model class manually.

Parameters:

Name Type Description Default
model_class Type[AbstractMotherPipeline]

The model class to register

required
algorithm Optional[str]

Optional algorithm name. If not provided, will be derived from class name

None

unregister_model(model_class_name)

Unregister a model class.

Parameters:

Name Type Description Default
model_class_name str

Name of the model class to unregister

required

algo_is_supported(algorithm)

Check if the specified algorithm is supported.

Parameters:

Name Type Description Default
algorithm str

Name of the algorithm to check

required

Returns:

Name Type Description
bool bool

True if supported, False otherwise

describe_model(name)

Get help text for a model class.

Parameters:

Name Type Description Default
name str

Name of the model class

required

Returns:

Name Type Description
str str

Help text for the model class

get_available_algorithms()

Get a list of all supported algorithms.

Returns:

Type Description
List[str]

List[str]: Names of supported algorithms

get_model_class(name)

Get a model class by name.

Parameters:

Name Type Description Default
name str

Name of the model class

required

Returns:

Name Type Description
Type type[AbstractMotherPipeline]

The model class

Raises:

Type Description
KeyError

If the model class is not found

get_model_class_by_algorithm(algorithm)

Get model classes by algorithm name.

Parameters:

Name Type Description Default
algorithm str

Name of the algorithm

required

Returns:

Name Type Description
Type list[Type[AbstractMotherPipeline]]

The model class

get_model_class_by_algorithm_and_type(algorithm, model_type)

Returns the appropriate model class based on the algorithm and model type.

get_supported_models()

Get a list of all supported model class names.

Returns:

Type Description
List[str]

List[str]: Names of supported model classes

register_model(algorithm=None)

Decorator to register a model class.

Parameters:

Name Type Description Default
algorithm Optional[str]

Optional algorithm name

None
Example

@register_model("my_algorithm") class MyCustomMother(AbstractMotherPipeline): pass

See the following example how to implement your custom RandomForest Classifier.

Example

Python
from mother import ml
from sklearn.ensemble import RandomForestClassifier

@ml.register_model("custom_rf")
class CustomRandomForestMother(RandomForestClassifier, ml.AbstractMotherPipeline):
    def get_hyperparameter_space(self, X, y, trial, prefix=""):
        return {
            f"{prefix}n_estimators": trial.suggest_int("n_estimators", 10, 100),
            f"{prefix}max_depth": trial.suggest_int("max_depth", 3, 10),
        }

    def default_parameters(self, prefix=""):
        return {f"{prefix}n_estimators": 50, f"{prefix}max_depth": 5}

print(ml.get_model_class_by_algorithm("custom_rf"))