pyfli.analysis.factor_analysis#
factor_analysis.py#
FactorAnalysis: bins pixel-wise FLI/FLIM results by a factor and compares how each fitting method’s estimated parameters vary across that factor.
Key idea#
decay, irf, and mask are the RAW / SHARED inputs that feed every fitting method in all_datasets / all_fitset – they are the same arrays regardless of which method produced a givenresult. Because of that, any factor computed directly from decay/irf (e.g. total_photons = decay.sum(time_axis)) is guaranteed to be pixel-for-pixel identical across methods – giving a true, common x-axis to compare methods against. This is different from (and safer than) binning by a method’s own estimated photon_count_map, which can differ in scale/definition between methods (that was the source of the earlier one-to-one mismatch).
Classes
|
- class FactorAnalysis(decay, irf, mask, all_datasets, all_fitset, method_names, time_axis=-1, sns_style='whitegrid', sns_palette='colorblind', cluster_mask=None, cluster_names=None, sns_cluster_palette='husl')[source]#
Bases:
object- Parameters:
decay (
ndarray) – Raw (binned) decay histogram, shape (…, T) with the time axis given by time_axis (default: last axis). All non-time axes together form the pixel grid, e.g. (H, W, T).irf (
ndarrayorNone) – Instrument response function, shared across all methods. Same time convention as decay. Can be per-pixel (matches decay’s pixel shape + time axis) or a smaller/1D IRF – both are accepted and simply stored as-is; only used directly by default factors/targets when its shape matches decay.mask (
ndarrayorNone) – Boolean mask shaped like decay’s pixel grid, shared across all methods. True = keep pixel. None disables masking.all_datasets (
list[dict]) – Per-method dictionaries of estimated parameter maps (e.g. {‘tau1_map’: …, ‘chi2_map’: …, …}), one dict per method, same order as method_names.all_fitset (
list[dict]) –- Per-method fit-result dicts. Each dict must contain at least:
- ’fit_map’reconstructed/predicted decay, same pixel grid +
time axis convention as decay, e.g. (H, W, T).
- ’residual_map’decay - fit_map (or the method’s own residual
definition), same shape as ‘fit_map’.
Extra keys (e.g. ‘sdf_map’, ‘convolved_map’) are allowed and may differ across methods – only ‘fit_map’/’residual_map’ are assumed to exist for every method. Use list_fitset_keys(i) / get_fitset_array(i, key) to work with any extra, method-specific keys. Same order as method_names.
method_names (
list[str]) – Label for each method, same order/length as all_datasets/all_fitset.time_axis (
int) – Axis of decay (and irf/fit_map/residual_map, when per-pixel) that indexes time bins. Default -1 (last axis).
- add_factor(name, array)[source]#
Register a pre-computed pixel-grid-shaped map as a reusable factor.
- add_factor_fn(name, fn)[source]#
Register a factor computed as fn(decay, irf) -> pixel-grid-shaped map.
- get_factor_map(factor_key)[source]#
- Resolve a factor map by name.
Shared factors derived from decay/irf (e.g. ‘total_photons’) – identical across all methods, so shared=True is returned.
A key present in every all_datasets[i] dict (a per-method estimated map). Use with caution: these are NOT guaranteed to be on a common scale/definition across methods, so shared=False.
- Returns:
maps (
list[ndarray] one map per method (same order as method_names))shared (
bool True if the SAME array object is used for every method)
- factor_values(factor_key, method_index=None)[source]#
Convenience accessor returning a single 2D factor map. - If factor_key is shared (e.g. ‘total_photons’), returns that one map
directly (no need to pick a method_index).
If factor_key is a per-method map (e.g. ‘fret_efficiency_map’), you must pass method_index to select which method’s version to use.
- cluster_selection_mask(cluster_id)[source]#
Boolean pixel-grid mask for one cluster label, intersected with the shared mask (if any).
- register_fitset_target(name, fn)[source]#
Add a derived target computed from a method’s fitset dict, e.g. a custom fit-quality metric. fn(decay, fitset_dict, irf) -> pixel-grid-shaped map, where fitset_dict is one entry of all_fitset (has ‘fit_map’, ‘residual_map’, and possibly extra method-specific keys – check with list_fitset_keys(method_index) before relying on anything beyond the two required keys).
- list_fitset_keys(method_index)[source]#
Keys actually present in all_fitset[method_index] for this method.
- get_fitset_array(method_index, key='fit_map')[source]#
Raw (H, W, T)-shaped array for one method’s fitset entry, e.g. fa.get_fitset_array(0, ‘fit_map’). Only ‘fit_map’ and ‘residual_map’ are guaranteed present for every method; other keys (e.g. ‘sdf_map’, ‘convolved_map’) may only exist for some methods – check with list_fitset_keys(method_index) first.
- selection_mask(factor_key, value_range, method_index=None)[source]#
Boolean pixel-grid mask(s) selecting pixels whose factor_key value falls within value_range=(low, high) (inclusive), AND passes the shared mask.
Returns a single 2D array if factor_key is shared (or method_index is given), otherwise a list of 2D arrays (one per method) since a per-method factor like ‘fret_efficiency_map’ selects different pixels per method.
- plot_spatial_selection(factor_key, value_range, ncols=3, figsize=None, cmap='viridis', bg_cmap='gray', saver=None, name=None)[source]#
Spatial map(s) of factor_key showing WHERE pixels fall inside value_range=(low, high). The full image is drawn in grayscale (bg_cmap) so the overall structure stays visible everywhere; ONLY the selected (in-range) pixels are drawn in color (cmap), on top.
- cmapcolormap for the selected (in-range) pixels – any matplotlib
colormap name, e.g. ‘viridis’, ‘plasma’, ‘magma’, ‘jet’.
bg_cmap : colormap for the grayscale structural background, default ‘gray’.
Produces ONE panel if factor_key is shared (same selection for every method), or one panel PER METHOD if it’s an individual per-dataset factor (e.g. ‘fret_efficiency_map’), since the selected pixels can differ by method.
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"spatial_selection_{factor_key}".
- plot_range_selection_grid(factor_key, value_range, target_keys, target_source='datasets', figsize=None, cmap='viridis', bg_cmap='gray', shared_scale=True, saver=None, name=None)[source]#
Combined grid: top row shows WHICH pixels are selected by factor_key in value_range=(low, high), one column per method; each subsequent row shows one target parameter map (e.g. ‘tau1_map’), restricted to that same selection. Every panel draws the full map in grayscale first (so the image’s structure is always visible), then overlays ONLY the in-range pixels in color on top – so you can visually compare how methods estimate a parameter for pixels drawn from a specific factor range (e.g. a photon-count band), without losing spatial context.
- cmapcolormap for the colored (in-range) overlay. Either:
a single string applied to every row, e.g. ‘viridis’, or
a list with one colormap PER ROW, ordered as [factor_key_row, target_keys[0]_row, target_keys[1]_row, …], e.g. cmap=[‘jet’, ‘plasma’, ‘plasma’] for factor_key=’total_photons’, target_keys=[‘tau1_map’, ‘tau2_map’]. Must have length 1 + len(target_keys).
bg_cmap : colormap for the grayscale structural background, default ‘gray’. shared_scale=True uses one colorbar range per target row (pooled across methods’ selected pixels) so panels are visually comparable; set False to let each panel auto-scale to its own selected pixels.
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"range_selection_grid_{factor_key}".
- analyze(factor_key='total_photons', target_keys=None, target_source='datasets', n_bins=10, bin_mode='quantile', bin_scope='auto', exclude_keys=('convergence_map', 'pixel_health_map'))[source]#
Bin pixels by factor_key and compute per-bin mean/median/std/sem/count of every target parameter, per method. The shared mask (if given) is applied for every method.
- factor_key can be:
a SHARED factor derived from decay/irf (e.g. ‘total_photons’) – identical values for every method, since decay/irf are common raw inputs. Example: fa.analyze(factor_key=’total_photons’, …).
an INDIVIDUAL per-method map already present in all_datasets (e.g. ‘fret_efficiency_map’) – each method computed its own version, so values/scales can legitimately differ by method. Example: fa.analyze(factor_key=’fret_efficiency_map’,
target_keys=[‘tau1_map’, ‘tau2_map’, ‘tau_mean_map’])
- target_source=’datasets’ pulls named maps from all_datasets[i][key]
(default target_keys = every key present).
- target_source=’fitset’ pulls derived maps computed from all_fitset’s
‘fit_map’/’residual_map’ vs. the shared decay (default target_keys = all registered fitset targets; see list_fitset_targets() / register_fitset_target()).
- bin_scope controls how bin edges are computed:
- ‘pooled’ONE set of edges from all methods’ values pooled
together. Correct when factor_key is shared (same array for every method) since edges then trivially agree with per-method edges too.
- ‘per_method’EACH method gets its own edges, computed only from
its own factor values. Use this when factor_key is an individual per-method map that may sit on a different scale/range per method (e.g. ‘fret_efficiency_map’) – pooling in that case can distort bins the way it did for mismatched photon-count scales. bin_rank (0..n_bins-1) is then the fair way to compare methods at “corresponding” positions along each one’s own distribution; the actual bin_center/bin_low/bin_high values may still differ by method.
- ‘auto’‘pooled’ if factor_key resolves to a shared factor,
‘per_method’ otherwise. This is a sensible default, not a strict rule – override explicitly if needed.
- Returns:
df (
pandas.DataFrame (long-form:method,bin_rank,bin_center,bin_low,) – bin_high, parameter, mean, median, std, sem, cv, count)bin_edges (
np.ndarrayordict[str,np.ndarray]) – A single edges array if bin_scope ended up ‘pooled’, otherwise a dict of {method_name: edges} for ‘per_method’.
- analyze_by_cluster(factor_key='total_photons', target_keys=None, method_index=0, target_source='datasets', n_bins=10, bin_mode='quantile', bin_scope='pooled', cluster_ids=None, exclude_keys=('convergence_map', 'pixel_health_map'))[source]#
Like analyze(), but bins/aggregates separately PER CLUSTER (from cluster_mask, given at construction) for a single method, instead of per method. Use this to see whether the factor/target relationship differs by spatial region (e.g. by ROI/letter), rather than by fitting method.
method_index selects which method’s datasets/fitset/per-method factor values to use (irrelevant when factor_key/target are shared).
- bin_scope‘pooled’ (default) uses ONE set of edges pooled across
all clusters – a common x-axis so clusters are directly comparable. ‘per_cluster’ gives each cluster its own edges (clusters then only comparable by bin_rank).
- Returns:
df (
pandas.DataFrame (long-form:cluster,bin_rank,bin_center,) – bin_low, bin_high, parameter, mean, median, std, sem, cv, count)bin_edges (
np.ndarrayordict[str,np.ndarray])
- plot(df, target_keys=None, kind='line', stat='mean', error='sem', ncols=3, figsize=None, logx=False, x_axis='auto', max_scatter_points=3000, scatter_alpha=0.35, scatter_size=10, box_showfliers=False, saver=None, name=None)[source]#
Grid of plots (one subplot per parameter, one series per method).
- kind‘line’, ‘box’, or ‘scatter’
- ‘line’ (default) – aggregated per-bin stat (‘mean’, ‘median’,
or ‘cv’ – coefficient of variation, std/mean) with a shaded error band (‘sem’, ‘std’, or None). Uses the pre-aggregated values already in df.
- ‘box’ – boxplot of the RAW per-pixel values within each
factor bin, one box per (bin, method) – shows spread and outliers instead of a single summary stat. box_showfliers controls whether outlier points are drawn.
- ‘scatter’ – raw per-pixel scatter of factor value (x, unbinned)
vs target value (y), one color per method – shows the actual relationship with no binning at all. Subsampled to max_scatter_points per (method, parameter) for plotting speed; control point look via scatter_alpha and scatter_size.
- x_axis (only used for kind=’line’)‘auto’, ‘bin_center’, or ‘bin_rank’
‘bin_center’ plots each method’s line at its own actual factor values – meaningful when bin_scope=’pooled’ (edges are identical across methods anyway). ‘bin_rank’ plots each method’s line at its bin index (0..n_bins-1) instead – use this when bin_scope=’per_method’, since each method’s bin_center values live on that method’s own scale and aren’t directly comparable at the same x position otherwise. ‘auto’ picks ‘bin_center’ if df.attrs[‘bin_scope’]==’pooled’, else ‘bin_rank’.
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"factor_analysis_{factor_key}_{kind}".
- plot_by_cluster(df, target_keys=None, stat='mean', error='sem', ncols=3, figsize=None, logx=False, x_axis='auto', saver=None, name=None)[source]#
Line-plot analog of plot(kind=’line’), for the output of analyze_by_cluster(): one subplot per parameter, one colored line per CLUSTER (using self.cluster_palette) instead of one per method.
stat : ‘mean’, ‘median’, or ‘cv’ (coefficient of variation, std/mean) error : ‘sem’, ‘std’, or None – shaded band around stat
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"factor_analysis_by_cluster_{factor_key}".
- static compare(labeled_dfs, target_keys=None, stat='mean', error='sem', ncols=3, figsize=None, logx=False, palette=None, saver=None, name=None)[source]#
Overlay analyze() results from INDEPENDENT FactorAnalysis instances (e.g. different simulated/acquired datasets, each with its own decay/irf/mask) onto shared subplots – one line per dataset, one subplot per parameter.
This is different from plot()/plot_by_cluster(), which compare methods/clusters WITHIN one shared decay: those assume every series was computed from the same raw decay, so a common factor_key like ‘total_photons’ is guaranteed pixel-for-pixel identical across series. compare() makes no such assumption – each df comes from its own instance’s own analyze() call, so each dataset’s line is drawn at its own actual bin_center x-values.
- labeled_dfslist[(str, pandas.DataFrame)]
(label, df) pairs, each df as returned by analyze(). All should use the same factor_key (and ideally similar binning) for the comparison to be meaningful.
- palettedict[str, color] or None
Maps each label to a color; defaults to a qualitative palette with one color per label.
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"factor_analysis_compare_{factor_key}".
- Return type:
fig,axes
- plot_bin_distribution(factor_key, target_key, target_source='datasets', n_bins=6, bin_mode='quantile', bin_scope='auto', kind='violin', color_by='auto', figsize=None, saver=None, name=None)[source]#
Seaborn violin/boxen plot of the RAW per-pixel distribution of target_key, split by factor bin (x-axis) – complements plot()’s aggregated mean/sem lines by showing the actual spread/shape of each bin’s pixel values, not just a summary stat.
kind : ‘violin’ or ‘boxen’ color_by : ‘auto’, ‘bin’, or ‘method’
‘method’ colors/dodges violins by method (needs the legend to tell methods apart). ‘bin’ gives each bin its own color instead (no per-method dodge). ‘auto’ picks ‘bin’ when there’s only one method (the common case – method color would be redundant with the x-axis) and ‘method’ otherwise.
- saverDataSaver-like object or None
If provided, the figure is saved via
saver.save_plot(name or default, fig=fig, close=False).- namestr or None
Explicit save name; defaults to
f"bin_distribution_{factor_key}_{target_key}".
- analyze_and_plot(factor_key='total_photons', target_keys=None, target_source='datasets', n_bins=10, bin_mode='quantile', bin_scope='auto', kind='line', stat='mean', error='sem', ncols=3, figsize=None, logx=False, x_axis='auto', saver=None, name=None)[source]#
Convenience wrapper: analyze() then plot() in one call.