Matchers

Propensity Score Matcher

class pybalance.propensity.PropensityScoreMatcher(matching_data, objective='beta', caliper=None, max_iter=50, time_limit=300, method='greedy', verbose=True, seed=None)[source]

Use a propensity score model to match two populations. The Matcher searches randomly over hyperparameters for the propensity score model and selects the match that performs best according to the given optimization objective.

Parameters:
  • matching_data (MatchingData) – Data containing pool and target populations to be matched.

  • objective (str | BaseBalanceCalculator) – Matching objective to optimize in hyperparameter search. Can be a string referring to any balance calculator known to utils.balance_calculators.BalanceCalculator or an instance of BaseBalanceCalculator.

  • caliper (float | None) – If defined, restricts matches to those patients with propensity scores within the caliper of each other. Note that caliper matching may lead to a loss of patients in the target population if no patient in the pool exists within the specified caliper. Should be in (0, 1].

  • max_iter (int) – Maximum number of hyperparameters to try before returning the best match.

  • time_limit (float | None) – Restrict hyperparameter search based on time. No new model will be trained after time_limit seconds have passed since matching began (def: 300 sec).

  • method (str) – Method to use for propensity score matching. Can be either ‘greedy’ or ‘linear_sum_assignment’. The former method is locally optimal and globally sub-optimial; the latter globally optimial but far more compute intensive. For large problems, use greedy.

  • verbose (bool) – Flag to indicate whether to print diagnositic information during training.

  • seed (int | None) – Seed for the hyperparameter search and the classifiers. Set it to make matching reproducible; if None, results vary between runs.

get_params()[source]

Return the matcher’s configuration parameters as a dict.

match()[source]

Match populations passed during __init__(). Returns MatchingData instance containing the matched pool and target populations.

Return type:

MatchingData

pybalance.propensity.plot_propensity_score_match_distributions(matcher)[source]

Plot histograms of the estimated propensity score for pool and target populations pre- and post-matching.

Parameters:

matcher (PropensityScoreMatcher) – Fitted PropensityScoreMatcher model.

pybalance.propensity.plot_propensity_score_match_pairs(matcher)[source]

Plot scatterplot of pool-target pairs formed by propensity score matching.

Parameters:

matcher (PropensityScoreMatcher) – Fitted PropensityScoreMatcher model.

Genetic Matcher

class pybalance.genetic.GeneticMatcher(matching_data, objective='beta', time_limit=300, verbose=True, candidate_population_size=None, n_candidate_populations=5000, n_keep_best=None, n_voting_populations=None, n_mutation=None, n_generations=1000, n_iter_no_change=100, max_batch_size_gb=2.0, seed=None, log_every=5)[source]

Match two populations using a genetic algorithm.

Parameters:
  • matching_data (MatchingData) – MatchingData to be matched. Must contain exactly two populations. The larger population will be matched to the smaller.

  • objective (str | BaseBalanceCalculator) – Matching objective to optimize. Can be a string referring to any balance calculator known to utils.balance_calculators.BalanceCalculator or an instance of BaseBalanceCalculator.

  • time_limit (float | None) – Time limit in seconds for matching. None means no limit. Defaults to 300 seconds.

  • verbose (bool) – Whether to print progress and diagnostic information.

  • candidate_population_size (int | None) – Size of candidate population to match against reference data. If not specified, will use the same size as the reference population.

  • n_candidate_populations (int) – Number of candidate populations to simultaneously evolve. Defaults to 5000.

  • n_keep_best (int | None) – Keep top N current best scoring candidate populations. If None, defaults to n_candidate_populations / 4.

  • n_voting_populations (int | None) – Form new candidate populations based on frequency of patient occurrence. If None, defaults to n_candidate_populations / 4.

  • n_mutation (int | None) – Make individual patient swaps for top candidate populations. If None, defaults to n_candidate_populations / 4.

  • n_generations (int) – Number of generations to evolve the candidate populations. Defaults to 1000.

  • n_iter_no_change (int) – Stop if no improvement after this many iterations. Defaults to 100.

  • max_batch_size_gb (float) – Maximum batch size in GB for GPU operations. Defaults to 2.0.

  • seed (int | None) – Random seed for reproducibility. If None, results will vary between runs.

  • log_every (int) – Log progress every N generations. Defaults to 5.

get_params()[source]

Return the matcher’s configuration parameters as a dict.

match()[source]

Match populations passed during __init__(). Returns MatchingData instance containing the matched pool and target populations.

Return type:

MatchingData

pybalance.genetic.get_global_defaults(n_candidate_populations=5000)[source]

Get a set of reasonable default values for evolutionary configuration. We break parameters into two groups: evolutionary, i.e., those that govern how the candidate populations are mixed, and initialization, i.e., those that govern the initial set of candidate populations.

Parameters:

n_candidate_populations – Number of candidate populations to evolve.

Constraint Satisfaction Matcher

class pybalance.lp.ConstraintSatisfactionMatcher(matching_data, objective='beta', pool_size=None, target_size=None, max_mismatch=None, time_limit=300, num_workers=4, ps_hinting=False, verbose=True)[source]

Population matching based on constraint satisfication formulation. This solver can only handle linear objective functions; see “objective” parameter below.

The constraints and optimization target are specified to the solver via the options pool_size, target_size, and max_mismatch. The behavior of the solver depends on which are these options are specified as given below:

(pool_size, target_size, max_mismatch) –> optimize balance subject to size and balance constraints

(pool_size, target_size) –> optimize balance subject to size constraints

(max_mismatch) –> optimize pool size subject to target_size = n_target and balance constraints

() –> optimize balance subject to size constraints with pool_size = target_size = n_target

Optimizing pool_size subject to balance constraint is known as “cardinality matching”. See https://kosukeimai.github.io/MatchIt/reference/method_cardinality.html and references therein.

Parameters:
  • matching_data (MatchingData) – A MatchingData object describing the pool and target populations. See utils.matching_data.

  • objective (str | BaseBalanceCalculator) – Matching objective to optimize. Technically, you can pass any balance calculator, but this solver cannot handle non-linear objective functions. The solver uses the preprocessing from the balance calculator for setting up the problem; the balance calculator itself is used to report the balance of generated matches but not in actually finding solutions (since the CS solver needs a discretized objective function). The solver will optimize the absolute mean difference on the output features of the balance calculator’s preprocessing.

  • pool_size (int | None) – Number of samples to include from the pool in the matched population. Must be less than the size of the pool. If pool_size is not set, then max_mismatch and target_size must be set and pool_size will be optimized subject to the target_size and max_mismatch constraints.

  • target_size (int | None) – Number of samples to include from the target in the matched population. Must be less than or equal to the size of the target. If target_size is not set,then max_mismatch and pool_size must be set and target_size will be optimized subject to the pool_size and max_mismatch constraints.

  • max_mismatch (float | None) – Maximum allowable absolute mean difference for any feature.

  • time_limit (float | None) – Time limit to stop solving in seconds (default 300 sec).

  • num_workers (int) – Number of workers to use to optimize objective. See https://github.com/google/or-tools/blob/stable/ortools/sat/sat_parameters.proto#L556 for more detail.

  • ps_hinting (bool) – Compute a propensity score match and use the result as a hint for the solver

  • verbose (bool) – Verbose solving.

get_params()[source]

Return the matcher’s configuration parameters as a dict.

match(hint=None)[source]

Match populations passed during __init__(). Returns MatchingData instance containing the matched pool and target populations.

Parameters:

hint (List[int] | None) –

You can supply a “hint” as either (1) A list of indices to the pool. It will be assumed that the entire target is used, or (2) A list of two lists, the first list being the indices to the target, and the second being the indices to the pool, or (3) by omitting the hint altoghether and passing ps_hinting=True in __init__(). In case (3), a propensity score model will be estimated on the fly and used to create a match population as a hint to the solver.

I admit the interface here is a bit confusing. We will clean this up in a later release.

Return type:

MatchingData

class pybalance.lp.AggregateConstraintSatisfactionMatcher(matching_data, pool_size=None, max_mismatch=None, time_limit=300, num_workers=4, verbose=True)[source]

Constraint-satisfaction matcher for an aggregate target: the pool is patient-level, but the reference population is known only through summary statistics (an AggregateTarget, e.g. a published Table 1). A subset of the pool is selected so that its moments match the target’s.

Every statistic the AggregateTarget discloses becomes a constraint, and nothing else is constrained: a feature’s mean (if disclosed), its variance (if a “std” is disclosed), each disclosed quantile / median / min / max (by dichotomizing the feature – see AggregateTargetBalanceCalculator) and each disclosed categoric rate. An undisclosed statistic or categoric level is simply left out. Setting a target std of 0 asks the solver for the minimal-variance subset achievable for that feature, subject to everything else. A std without a mean is the variance around the subset’s own mean.

The target is a fixed constraint vector and is never subsetted, so several ConstraintSatisfactionMatcher options do not apply here and are intentionally absent:

  • no objective: only the beta / mean-matching formulation (via AggregateTargetBalanceCalculator) is defined against aggregate moments, so it is used unconditionally;

  • no target_size / match_size: the target cannot be subsetted, so the effective target size is always AggregateTarget.n;

  • no ps_hinting / solver hints: there is no patient-level target to build a hint from.

Parameters:
  • matching_data (MatchingData) – A MatchingData object whose target is an AggregateTarget. See utils.matching_data.

  • pool_size (int | None) – Number of samples to include from the pool in the matched population. Must be less than the size of the pool. If neither pool_size nor max_mismatch is set, pool_size defaults to AggregateTarget.n. If max_mismatch is set and pool_size is not, then pool_size is optimized subject to the balance constraint (cardinality matching).

  • max_mismatch (float | None) – Maximum allowable absolute mean difference for any feature. Only applies to mean matching; variance matching (for features with a disclosed target std) is always a soft, minimized objective term and is not capped by max_mismatch.

  • time_limit (float | None) – Time limit to stop solving in seconds (def: 300 sec).

  • num_workers (int) – Number of workers to use to optimize objective.

  • verbose (bool) – Verbose solving.

get_params()[source]

Return the matcher’s configuration parameters as a dict.

match()[source]

Match the pool to the aggregate target passed during __init__(). Returns a MatchingData instance containing the matched pool and the (unchanged) aggregate target.

Return type:

MatchingData