Core Utilities
Matching Data
- class pybalance.utils.MatchingHeaders(categoric, numeric)[source]
MatchingHeaders is a simple data structure to store information about which features to be used for matching and separating these features into categoric (e.g. country, gender) and numeric (e.g. age, weight) types.
- Parameters:
categoric (List[str]) – List of features to be treated as categoric variables.
numeric (List[str]) – List of features to be treated as numeric variables.
- class pybalance.utils.MatchingData(data=None, headers=None, population_col='population', pool=None, target=None, pool_name='pool', target_name='target')[source]
It is common in matching problems to require basic metadata about the data in order to perform matching. For instance, the data may contain columns such as “patient_id”, “population” and “index_date”, which are not intended to be used for matching but which must “go along for the ride” and follow the main data everywhere. MatchingData is a wrapper around pandas.DataFrame that includes this additional required logic about the columns. Features required for matching are described by a “headers” field, while other columns exist alongside. See MatchingHeaders.
Construction patterns:
# Combined patient-level table with a population column MatchingData(df) MatchingData(data=df, population_col="population") # Explicit patient-level pool and target MatchingData(pool=pool_df, target=target_df) # Patient-level pool and aggregate-only target MatchingData(pool=pool_df, target=AggregateTarget.from_dict(...))
- Parameters:
data (Optional[Union[pd.DataFrame, str]]) – Data frame containing both matching feature data for all populations as well as at least one additional column specifying to which population each row belongs. If a string is passed, it is assumed to be a path to the data frame. Mutually exclusive with
pool/target.headers (Optional[MatchingHeaders]) – A MatchingHeaders object with keys “numeric” and “categoric” and whoses values are names of columns to be used for matching. If None is passed, headers will be inferred based on how many unique values each column has (or from
AggregateTargetwhen the target is aggregate). As guessing the headers can lead to errors, it is recommended to supply them explicitly.population_col (str) – Name of the column used to split data into subpopulations.
pool (Optional[Union[pd.DataFrame, str]]) – Patient-level pool population. Used with
target.target (Optional[Union[pd.DataFrame, str, AggregateTarget]]) – Either a patient-level target DataFrame or an
AggregateTargetsummary. Used withpool.pool_name (str) – Population label for
poolwhen using the explicit split constructor.target_name (str) – Population label for
targetwhen using the explicit split constructor.
- append(df, name=None)[source]
Append a population to an existing MatchingData instance. This operation is inplace.
- property data: DataFrame
Pointer to underlying pandas DataFrame.
- describe(normalize=True, aggregations=['mean', 'std'], quantiles=[0, 0.25, 0.5, 0.75, 1])[source]
Calls describe_categoric() and describe_numeric() and returns the results in a single dataframe.
- Return type:
DataFrame
- describe_categoric(normalize=True)[source]
Create a summary statistics table split by population for categoric variables.
- Return type:
DataFrame
- describe_numeric(aggregations=['mean', 'std'], quantiles=[0, 0.25, 0.5, 0.75, 1], long_format=True)[source]
Create a summary statistics table split by population for numeric variables.
- Return type:
DataFrame
- get_population(population)[source]
Get the matching data for a population by its name.
- Return type:
DataFrame
- property populations: List[str]
List of all populations present in the MatchingData object.
- pybalance.utils.infer_matching_headers(data, max_categories=10, ignore_cols=['patient_id', 'patientid', 'population', 'index_date'])[source]
This utility function guesses which columns are numeric and which columns are categoric from input data. The data can be passed either as separate data frames target and pool or combined in one and passed with keyword argument data. The function returns a dictionary with keys ‘numeric’, ‘categoric’ and ‘all’ with values equal to the list of column names of the given type. By default, the function ignores patient_id and population columns.
- Return type:
- pybalance.utils.split_target_pool(matching_data, pool_name=None, target_name=None)[source]
Split matching_data into target and pool populations based. If the names of the target and pool populations are not explicitly provided, the routine will attempt to infer their names, assuming that the target population is the smaller population.
- Return type:
DataFrame
Preprocessing
- class pybalance.utils.BaseMatchingPreprocessor[source]
BaseMatchingPreprocessor is an abstract preprocessor class for organizing data transformations on matching data, keeping track of all preprocessing steps so that data are always transformed transformed in the same way.
The inherited class must implement:
_fit(), _transform(), _get_output_headers()
and should implement, if possible,
_get_feature_names_out()
The class extends the preprocessing classes defined in sklearn. In addition to handling transformations of the data, this class also handles logic of MatchingHeaders. Preprocessors which conform to the BaseMatchingPreprocessor standard are chainable, allowing one to easily combine preprocessing tasks; see ChainPreprocessor.
- abstract _fit(matching_data)[source]
This method performs all calculations needed in order to perform the transformation tasks of the preprocessor (e.g. computing means and standard deviations). Should accept a MatchingData instance as input and return None. This method must be overridden by the subclass.
- _get_feature_names_out(feature_name_in)[source]
Same as get_feature_names_out but on a single feature level. This method should be overridden by the subclass.
- Return type:
List[str]
- abstract _get_output_headers()[source]
Return headers on the output matching data. This method must be overridden by the subclass.
- class pybalance.utils.CategoricOneHotEncoder(drop='first')[source]
CategoricOneHotEncoder converts categoric covariates into one-hot encoded variables. Numeric columns are unaffected.
- Parameters:
drop (str | None) – Which, if any, columns to drop in the transformation. Choices are: {‘first’, ‘if_binary’, None}. See sklearn.preprocessing.OneHotEncoder for more details.
- class pybalance.utils.NumericBinsEncoder(n_bins=5, strategy='uniform', encode='onehot-dense', cumulative=False)[source]
NumericBinsEncoder discretizes numeric covariates according to specified binning strategy. Categoric columns are unaffected.
- Parameters:
n_bins (int) – Number of bins to split numeric variable into. Note in the case of cumulative = True, the last bin would be always one. To avoid this, internally we use n_bins + 1 bins and drop the last bin.
strategy (str) – Strategy to use for binnings. Choices are: {‘uniform’, ‘quantile’, ‘kmeans’}. See sklearn.preprocessing.KBinsDiscretizer for more details.
encode (str) – Method to use for encoding numeric variable. Choices are: {‘onehot’, ‘onehot-dense’, ‘ordinal’}. See sklearn.preprocessing.KBinsDiscretizer for more details.
cumulative (bool) – Whether to transform numeric variables to discretized cumulative distribution. E.g. if x is in the 3rd bin and n_bins = 4, then x will map to [0, 0, 1, 0] if cumulative = False or [0, 0, 1, 1] otherwise.
- class pybalance.utils.DecisionTreeEncoder(keep_original_features=False, **decision_tree_params)[source]
DecisionTreeEncoder transforms all covariates into binary coviates corresponding to their terminal (leaf) position on a decision tree.
- class pybalance.utils.ChainPreprocessor(preprocessors)[source]
ChainPreprocessor applies a sequences of Preprocessors.
- Parameters:
preprocessors (List[BaseMatchingPreprocessor]) – A list of preprocessors to be applied in sequence to the input data.
Balance Calculators
- class pybalance.utils.BaseBalanceCalculator(matching_data, preprocessor, feature_weights=None, order=1, standardize_difference=True, device=None)[source]
BaseBalanceCalculator is the low-level interface to calculating balance. BaseBalanceCalculator can be used with any preprocessor defined as a subclass of BaseMatchingPreprocessor. BaseBalanceCalculator implements matrix calculations in pytorch to allow for GPU acceleration.
BaseBalanceCalculator performs two main tasks:
Computes a per-feature-loss based on the output features of the given preprocessor and
Aggregates the per-feature-loss into a single value for the loss.
Furthermore, the calculator can compute the loss for many populations at a time.
- Matching_data:
Input matching data to be used for distance calculations. Must contain exactly two populations. The smaller population is used as a reference population. Calls to distance() compute the distance to this reference population.
- Preprocessor:
Preprocessor to use for per-feature-loss calculation. The per-feature-loss is, up to some normalizations, the mean difference in the features at the output of the preprocessor.
- Feature_weights:
How to weight features in aggregation of per-feature-loss.
- Order:
Exponent to use in combining per-feature-loss into an aggregate loss. Total loss is sum(feature_weight * feature_loss**order)**(1/order).
- Parameters:
standardize_difference (bool) – Whether to use the absolute standardized mean difference for the per-feature loss (otherwise uses absolute mean difference).
- Device:
Name of device to use for matrix computations. By default, will use GPU if a GPU is found on the system.
- class pybalance.utils.BetaBalance(matching_data, feature_weights=None, device=None, drop='first', standardize_difference=True)[source]
BetaBalance computes the balance between two populations as the mean absolute standardized mean difference across all features. Uses StandardMatchingPreprocessor as the preprocessor. In this preprocessor, numeric variables are left unchanged, while categorical variables are one-hot encoded. See StandardMatchingPreprocessor for more details.
- class pybalance.utils.BetaSquaredBalance(matching_data, feature_weights=None, device=None, drop='first', standardize_difference=True)[source]
BetaSquaredBalance computes the balance between two populatiosn as the mean square standardized mean difference across all features. Uses StandardMatchingPreprocessor as the preprocessor. In this preprocessor, numeric variables are left unchanged, while categorical variables are one-hot encoded. See StandardMatchingPreprocessor for more details.
- class pybalance.utils.BetaMaxBalance(matching_data, feature_weights=None, device=None, drop='first', standardize_difference=True)[source]
Same as BetaBalance, except the worst-matched feature determines the loss. This class is provided as a convenience, since this balance metric is often a criterion used to determine if matching is “sufficiently good”. However, be aware that using this balance metric as an optimization objective with the various matchers can lead unwanted behavior, since if improvements in the worst-matched feature are not possible, there is no signal from the balance function to improve any other the other features.
- class pybalance.utils.GammaBalance(matching_data, feature_weights=None, device=None, n_bins=5, encode='onehot-dense', cumulative=True, drop='first', standardize_difference=True)[source]
Convenience interface to BaseBalanceCalculator to compute the balance between two populations by computing the mean area between their one-dimensional marginal distributions. See GammaPreprocessor for description of preprocessing options.
- class pybalance.utils.GammaSquaredBalance(matching_data, feature_weights=None, device=None, n_bins=5, encode='onehot-dense', cumulative=True, drop='first', standardize_difference=True)[source]
Same as GammaBalance, except that per-feature balances are averages in a mean square fashion.
- class pybalance.utils.GammaXTreeBalance(matching_data, keep_original_features=False, device=None, standardize_difference=True, **decision_tree_params)[source]
- pybalance.utils.BalanceCalculator(matching_data, objective='gamma', **kwargs)[source]
BalanceCalculator provides a convenience interface to balance calculators, allowing the user to initialize a balance calculator by name. The calculators are initialized with default parameters, but these can be overridden by passing the appropriate kwargs.
- Parameters:
matching_data – MatchingData instance containing reference to the data against which matching metrics will be computed
objective – Name of objective function to be used for computing balance. Balance calculators must be implemented in utils.balance_calculators.py and registered in the BALANCE_CALCULATORS dictionary therein in order to be accessible from this interface.
kwargs – Any additional arguments required to configure the specific objective function (e.g. n_bins = 10 for “gamma”).