Visualization
- pybalance.visualization.plot_numeric_features(matching_data, col_wrap=2, height=6, include_only=None, **plot_params)[source]
Plot the one-dimensional marginal distributions for all numerical features and all treatment groups found in matching_data. Extra keyword arguments are passed to seaborn.histplot and override defaults.
If matching_data has an AggregateTarget, it has no patient-level rows to plot a distribution for and a warning is logged; use plot_aggregate_target_match() instead to compare the pool against its disclosed statistics.
- Parameters:
matching_data (MatchingData) – MatchingData instance containing at least one population.
include_only (List[str] | None) – List of features to consider for plotting. Otherwise, all numeric features are plotted.
- Return type:
Figure
- pybalance.visualization.plot_categoric_features(matching_data, col_wrap=2, height=6, include_binary=True, include_only=None, **plot_params)[source]
Plot the one-dimensional marginal distributions for all categoric features and all treatment groups found in matching_data. Extra keyword arguments are passed to seaborn.histplot and override defaults.
If matching_data has an AggregateTarget, it has no patient-level rows to plot a distribution for and a warning is logged; use plot_aggregate_target_match() instead to compare the pool against its disclosed statistics.
- Parameters:
matching_data (MatchingData) – MatchingData instance containing at least one population.
include_binary – Whether to include binary features in the plot.
include_only (List[str] | None) – List of features to consider for plotting. Otherwise, all categoric features are plotted. If include_binary is False, binary features are excluded, even if present in include_only.
- Return type:
Figure
- pybalance.visualization.plot_binary_features(matching_data, max_features=25, include_only=None, orient_horizontal=False, standardize_difference=False, reference_population=None, **plot_params)[source]
Plot all binary features for all treatment groups found in matching_data. Additional keyword arguments are passed to sns.barplot and override default.
- Parameters:
matching_data (MatchingData) – MatchingData instance containing at least a pool and target population.
max_features (int) – Max number of features to show in plot, in case there are a lot of binary features. Features are sorted in descending order by the initial mismatch between pool and target. The top max_features will be shown.
include_only (List[str] | None) – List of features to consider for plotting. Otherwise, all binary features are plotted.
orient_horizontal (bool) – If True, orient features along the x-axis. Otherwise, features will be along the y-axis.
standardize_difference (bool) – Whether to use the absolute standardized mean difference for the differences plot (otherwise plots absolute mean difference).
reference_population (str | None) – Name of population in matching_data against which other populations should be compared. If not supplied, will use the smaller population as the reference population.
plot_params – Parameters passed on to seaborn routines.
- Return type:
Figure
- pybalance.visualization.plot_joint_numeric_distributions(matching_data, joint_kind='kde', include_only=None, **plot_params)[source]
Plot 2D distributions of pairs of numeric features from matching_data. joint_kind can be either kde or scatter. scatter is usually a bad choice for large datasets. Choose subsets of features using include_only. Additional keyword arguments are passed to sns.JointGrid and override default.
- pybalance.visualization.plot_joint_numeric_categoric_distributions(matching_data, include_only_numeric=None, include_only_categoric=None, **plot_params)[source]
Plot 2D distributions of pairs of numeric and categoric features from matching_data. Choose subsets of features using include_only. Additional keyword arguments are passed to sns.JointGrid and override default.
- pybalance.visualization.plot_per_feature_loss(matching_data, balance_calculator, reference_population=None, debin=True, normalize=False, **plot_params)[source]
Plot the mismatch as a function of feature.
- Parameters:
matching_data (MatchingData) – Input data to plot.
balance_calculator (BaseBalanceCalculator) – Balance metric to use for calculating the per feature loss. Balance calculator must implement a ‘per_feature_loss’ method.
reference_population (str | None) – Name of population in matching_data against which other populations should be compared. If not supplied, will use the smaller population as the reference population.
debin (bool) – If True, attempt to map effective features back into the real feature space. This is not always possible, e.g., features like age*height can’t be mapped back to a single feature but features like country_US, country_Germany can. In the former case, routine will plot loss per effective feature; in the latter, loss per input feature.
normalize (bool) – If True, divide loss by number of features such that the sum is the total loss. Otherwise, the plotted loss contributions must be averaged to obtain the total loss.
plot_params – Parameters passed on to seaborn routines.
- Return type:
Figure