Weighters

Entropy Balancing & MAIC

class pybalance.weighting.BaseWeighter(matching_data, weight_col='sample_weight', verbose=True)[source]

Common interface for weighting methods. Unlike the Matcher classes (genetic/lp/propensity), a Weighter never drops pool patients; instead it assigns every pool patient a non-negative weight so that the weighted pool resembles the target on the moments of interest. This is the standard approach for indirect/external comparisons when excluding patients is undesirable or infeasible – e.g. small pools, or a target known only through published aggregate statistics.

Parameters:
  • matching_data (MatchingData) – MatchingData whose pool is to be weighted. The target can be either patient-level or an AggregateTarget (e.g. a published Table 1).

  • weight_col (str) – Name of the column used to store weights on the MatchingData returned by match(). Must not collide with an existing matching feature.

  • verbose (bool) – Whether to log fitting diagnostics.

effective_sample_size()[source]

Kish’s effective sample size of the fitted weights. See effective_sample_size().

Return type:

float

get_params()[source]

Return the weighter’s configuration parameters as a dict.

Return type:

Dict

get_weights()[source]

Return the fitted per-patient pool weights (in pool row order).

Return type:

ndarray

match()[source]

Fit (if not already fit) and return a MatchingData instance in which every pool patient is retained and carries a new column (see weight_col) holding their fitted weight. The target population (patient-level or aggregate) is passed through unchanged, with weight 1.0 assigned to any patient-level target rows.

Return type:

MatchingData

class pybalance.weighting.EntropyBalanceWeighter(matching_data, match_variance=False, normalize='target', max_iter=200, tol=1e-08, ridge=1e-08, weight_col='sample_weight', verbose=True, limit_penalty=0.01)[source]

General maximum-entropy (“method of moments”) weighting: solves for pool weights of the form w_i = exp(z_i . alpha) such that the weighted pool matches the target. For an AggregateTarget every disclosed statistic is its own constraint and nothing else is constrained: a numeric mean, the rate of each disclosed categoric level, and the rate above each disclosed median / quantile (the feature is dichotomized at the disclosed value). A feature or categoric level the target says nothing about is simply left free. Optionally also balances variance for numeric features whose target discloses (or, for a patient-level target, has) a standard deviation; for an AggregateTarget that needs the mean too, since the variance is taken around it.

A disclosed min / max is a soft constraint, as in AggregateConstraintSatisfactionMatcher: exactly zero weight on patients above a max cannot be reached by positive weights, so the fraction of weight there is only penalized (see limit_penalty).

This is the general form of Matching-Adjusted Indirect Comparison (MAIC; Signorovitch et al., 2010); see MAICWeighter for the classic mean-only formulation.

Parameters:
  • matching_data (MatchingData) – MatchingData whose pool is to be weighted.

  • match_variance (bool) – If True, additionally constrain the weighted variance of numeric features. For an AggregateTarget, only features that actually disclose a “std” are constrained (mirroring AggregateConstraintSatisfactionMatcher); for a patient-level target, every numeric feature’s variance is constrained. Categoric features are never variance-constrained: their variance is already determined by their (matched) rate.

  • normalize (str) – How to rescale the fitted weights for reporting: “target” (default) rescales so weights sum to the target population size, “pool” rescales so weights sum to the pool size, “none” leaves the raw dual-optimizer weights unscaled. This choice has no effect on balance (the moment constraints are scale-invariant) or on effective_sample_size(); it only affects the units the weights are reported in.

  • max_iter (int) – Maximum number of Newton iterations.

  • tol (float) – Convergence tolerance on the (max-norm) constraint violation.

  • ridge (float) – Ridge regularization added to the Newton step for numerical stability; increase this if fitting fails to converge due to collinear/near-collinear features.

  • weight_col (str) – Name of the column used to store weights on the MatchingData returned by match().

  • verbose (bool) – Whether to log fitting diagnostics.

  • limit_penalty (float) – Strength of the quadratic penalty on the weight placed beyond a disclosed min / max of an AggregateTarget (a soft constraint). Larger values enforce the limit more tightly.

get_params()[source]

Return the weighter’s configuration parameters as a dict.

Return type:

Dict

class pybalance.weighting.MAICWeighter(matching_data, normalize='target', max_iter=200, tol=1e-08, ridge=1e-08, weight_col='sample_weight', verbose=True, limit_penalty=0.01)[source]

Matching-Adjusted Indirect Comparison (MAIC; Signorovitch et al., 2010, “Comparative effectiveness without head-to-head trials: a method for matching-adjusted indirect comparisons applied to psoriasis clinical trials”). Reweights patient-level (“IPD”) pool data so that its weighted means match a target’s aggregate statistics – typically a comparator trial’s published Table 1 – using the method-of-moments / maximum-entropy weighting scheme of EntropyBalanceWeighter, restricted to first moments only. This is the standard MAIC formulation and is appropriate whenever the comparator discloses only means (and category rates), not variances.

Use EntropyBalanceWeighter(matching_data, match_variance=True) directly if the comparator additionally discloses standard deviations you also want to match.

Parameters:
  • matching_data (MatchingData) – MatchingData whose pool (IPD) is to be weighted to match the target’s (typically aggregate, e.g. AggregateTarget) moments.

  • normalize (str) – See EntropyBalanceWeighter. Defaults to “target”, i.e. weights sum to the target’s sample size, matching common MAIC reporting conventions.

  • max_iter (int) – Maximum number of Newton iterations.

  • tol (float) – Convergence tolerance on the (max-norm) constraint violation.

  • ridge (float) – Ridge regularization added to the Newton step for numerical stability.

  • weight_col (str) – Name of the column used to store weights on the MatchingData returned by match().

  • verbose (bool) – Whether to log fitting diagnostics.

  • limit_penalty (float) – See EntropyBalanceWeighter.

get_params()[source]

Return the weighter’s configuration parameters as a dict.

Return type:

Dict

IPTW Weighter

class pybalance.weighting.IPTWWeighter(matching_data, classifier=None, trim_quantiles=None, weight_col='sample_weight', verbose=True)[source]

Inverse Probability of Treatment Weighting (IPTW; Rosenbaum & Rubin, 1983; see Austin, 2011, “An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies”, for the ATT construction used here). Fits a propensity model p(X) = P(target | X) that classifies pool vs. target patients on their covariates, then reweights each pool patient by the odds p / (1 - p). The target population keeps weight 1, so the weighted pool is reweighted onto the target’s covariate distribution – i.e. this is the ATT estimand with the target playing the role of the fixed/reference (“treated”) group, which matches the convention used by MAICWeighter/EntropyBalanceWeighter and the package’s typical use case of building an external comparator arm that represents a trial population.

Unlike EntropyBalanceWeighter, IPTW does not guarantee exact balance on any particular moment: it only guarantees balance asymptotically, and only if the propensity model is correctly specified. It also requires a patient-level target (there is no pool-vs-target classification problem to fit against a published aggregate Table 1). Its main practical advantages over entropy balancing are that it scales to many covariates without requiring the target’s moments to be inside the pool’s convex hull, and that it is the most widely recognized/reported method in the observational literature.

Parameters:
  • matching_data (MatchingData) – MatchingData whose pool is to be weighted. The target must be patient-level (not an AggregateTarget).

  • classifier (Optional[BaseEstimator]) – A fitted-or-unfitted sklearn-compatible classifier exposing predict_proba, used to estimate P(target | X). It is cloned before fitting, so passing a pre-fitted instance does not reuse its fit. Defaults to LogisticRegression(max_iter=1000), the standard choice for propensity score estimation.

  • trim_quantiles (Optional[Tuple[float, float]]) – Optional (low, high) quantiles (e.g. (0.01, 0.99)) at which to clip the fitted weights. Extreme weights (driven by pool patients whose covariates make them look almost certainly pool or almost certainly target) are the main practical failure mode of IPTW; trimming trades a little bias for a large reduction in variance. Left unset (None) by default so the raw weights are returned untouched.

  • weight_col (str) – Name of the column used to store weights on the MatchingData returned by match().

  • verbose (bool) – Whether to log fitting diagnostics.

get_params()[source]

Return the weighter’s configuration parameters as a dict.

Return type:

Dict

pybalance.weighting.plot_iptw_propensity_distributions(weighter)[source]

Plot histograms of the estimated propensity score for the pool and target populations, before vs. after IPTW weighting – the weighting analogue of pybalance.propensity.plot_propensity_score_match_distributions.

Unlike a Matcher, a Weighter never drops patients, so there is no matched subset to compare against; instead, the “before” panel shows every pool patient counted equally and the “after” panel shows the same propensity scores counted by their fitted IPTW weight (the target always keeps weight 1). A successful fit should show the “after” pool histogram move towards the target’s.

Parameters:

weighter (IPTWWeighter) – A fitted IPTWWeighter, i.e. one on which match() has already been called.

Shared Utilities

pybalance.weighting.effective_sample_size(weights)[source]

Kish’s effective sample size: (sum w)**2 / sum(w**2). Invariant to the overall scale of the weights. A large drop relative to len(weights) indicates the reweighting is relying heavily on a small number of pool patients to match the target, and results should be interpreted cautiously (e.g. the target may lie outside the range of the pool’s covariates).

Return type:

float

pybalance.weighting.weighted_balance_table(weighter)[source]

Return a table comparing each balanced feature’s weighted (and, for reference, unweighted) pool moment to the target moment – a quick post-hoc check of how well match() balanced the pool. For EntropyBalanceWeighter/MAICWeighter, residuals should be ~0 for every constrained row; for IPTWWeighter, which does not solve for exact balance, this is instead a diagnostic of how much balance improved relative to the unweighted pool.

Parameters:

weighter (BaseWeighter) – A fitted EntropyBalanceWeighter/MAICWeighter/IPTWWeighter, i.e. one on which match() has already been called.

Return type:

DataFrame