All tests, models, and statistical analyses.
Descriptives & basic tests
Descriptives
What it is. Mean, median, SD, min, max, quartiles, skewness, kurtosis.
When to use. First thing to run on any numeric variable.
Frequencies
What it is. Frequency table: counts, %, and cumulative % per value.
When to use. For categorical or discrete variables.
T-Test
What it is. Compare means. One-sample (vs value), Independent (two groups), Paired (repeated measures).
Options / Parameters
Levene's test — Checks equality of variances (important for independent-samples).Cohen's d — Effect size — small 0.2, medium 0.5, large 0.8.
ANOVA
What it is. Compare means across 3+ groups.
Options / Parameters
One-Way — Single categorical factor.Two-Way — Two crossed factors — tests main effects + interaction.Repeated Measures — Same subjects measured under several conditions.Post-hoc Tukey HSD — Pairwise comparisons if the ANOVA is significant.Levene — Homogeneity of variance (assumption check).eta² — Proportion of variance explained by the factor.
Chi-Square
What it is. Tests association between two categoricals, or goodness-of-fit to an expected distribution.
Options / Parameters
Independence — Two categorical vars — produces crosstab + Cramer's V.Goodness of Fit — One variable — tests if observed freqs match expected (uniform or custom).
Correlation
What it is. Linear (Pearson), monotonic (Spearman), or rank-based (Kendall) association.
Options / Parameters
Pearson — Linear. Requires normality and interval scale.Spearman — Robust to outliers; measures monotonic relationships.Kendall — Better for small samples, more conservative.Partial — Correlation controlling for a third variable.
Regression
Linear Regression and Variants
What it is. Model a numeric outcome as a function of predictors.
Options / Parameters
OLS — Standard — minimizes sum of squared residuals.Ridge (L2) — Regularized — shrinks coefficients; helps with multicollinearity.LASSO (L1) — Can drive coefficients to zero — automatic feature selection.Elastic Net — L1+L2; alpha controls strength, l1_ratio mixes the penalties.Robust (Huber) — Resistant to outliers.Quantile — Estimates quantiles (e.g., median) instead of the mean.Stepwise — Adds/removes predictors automatically by AIC/BIC.WLS — Weighted LS — each observation has a weight (pick a weight variable).FGLS — Feasible Generalized LS — corrects heteroscedasticity/serial correlation.IV-2SLS — Instrumental variables — endogenous predictors replaced via instruments.ARDL — Autoregressive Distributed Lag — short- and long-run effects.Truncated — Y bounded (e.g., Y≥0).NLS — Nonlinear LS — provide a formula and starting values (JSON).Robust SE (HC3) — Heteroscedasticity-consistent standard errors.Standardized β — Coefficients in SD units — compare relative importance.
For a binary outcome, use Logistic (dedicated menu entry), not OLS.
Logistic Regression
What it is. Binary (0/1) outcome modeled as a probability.
Options / Parameters
Odds ratios — exp(β) — odds-ratio interpretation.McFadden R² — Pseudo-R² equivalent.Classification accuracy — How many cases are correctly classified.
Count Models
What it is. For Y that is a count (non-negative integer).
Options / Parameters
Poisson — Assumes E(Y)=Var(Y).Negative Binomial — For data with overdispersion (Var > E).ZIP — Zero-inflated Poisson — mixes structural zeros with counts.ZINB — Like ZIP but with overdispersion.Hurdle — Binary (zero/non-zero) + truncated count model.
Non-Parametric
Mann-Whitney U
What it is. Alternative to independent-samples t-test when normality fails.
Wilcoxon signed-rank
What it is. Alternative to paired t-test.
Kruskal-Wallis
What it is. Alternative to one-way ANOVA.
Non-Parametric Extended
What it is. Bundle of additional distribution-free tests.
Options / Parameters
KS 2-sample — Do two samples come from the same distribution?Sign test — Simple alternative to paired t-test.Runs test — Tests randomness of a sequence.McNemar — Paired nominal (before/after).Cochran Q — Multiple related samples with binary outcomes.Kendall's W — Agreement between raters/rankings.Jonckheere-Terpstra — Ordered alternative to Kruskal-Wallis.Mood's median — Tests equality of medians across k groups.Permutation — Exact resampling-based tests.Panel Unit Root — Stationarity for panel data.KDE — Kernel density estimation.LOESS — Local regression / smoothing.
Dimension reduction & reliability
PCA / EFA / Canonical Correlation
What it is. Dimension reduction and latent structure.
Options / Parameters
PCA — Orthogonal components maximizing variance. Use Kaiser (eigenvalue>1) to pick count.EFA — Factor analysis with rotation (varimax/promax/oblimin). Includes KMO + Bartlett.Canonical Corr. — Correlations between TWO sets of variables.
Clustering
What it is. Group similar observations.
Options / Parameters
K-Means — K-cluster partition by centroids. Pick K via silhouette score.Hierarchical — Agglomerative — produces a dendrogram. Linkage: ward / complete / average / single.
Reliability
What it is. Internal consistency of a scale.
Options / Parameters
Cronbach α — α ≥ 0.9 excellent, ≥ 0.8 good, ≥ 0.7 acceptable.Split-Half — Splits the scale in halves, correlates, applies Spearman-Brown correction.
Time Series
Unit Root Tests
What it is. Tests whether a series is stationary.
Options / Parameters
ADF — H0 = unit root (non-stationary).KPSS — Opposite null — stationarity.Phillips-Perron — Non-parametric correction for serial correlation.
Run ADF + KPSS together for a robust conclusion.
ARIMA
What it is. Autoregressive Integrated Moving Average model.
Options / Parameters
Auto-select (p,d,q) — Grid search minimizing AIC.p, d, q (manual) — AR lags, differencing, MA lags.Combine Forecasts — Combine forecasts from multiple models (mean/weighted).
GARCH family
What it is. Models time-varying volatility (volatility clustering).
Options / Parameters
GARCH(p,q) — Standard symmetric.EGARCH — Exponential — captures asymmetry (leverage effect).GJR-GARCH — Threshold — bad news increases volatility more.DCC-GARCH — Multivariate — dynamic conditional correlations across series.TAR — Threshold AR — different regimes above/below a threshold.
VAR & Granger Causality
What it is. Model multiple series jointly.
Options / Parameters
VAR — Vector autoregression with auto lag selection.Granger test — X 'causes' Y if lags of X improve Y's forecast.Cointegration — Individually non-stationary series whose linear combo is.
Panel & Causal
Panel Data (FE/RE/Hausman)
What it is. Longitudinal data (multiple entities × time).
Options / Parameters
Fixed Effects — Removes time-invariant unobserved heterogeneity.Random Effects — Assumes effects uncorrelated with X — more efficient.Hausman test — Decides FE vs RE. Rejecting H0 → use FE.
Causal Inference
What it is. Causal identification via quasi-experimental designs.
Options / Parameters
DiD — Difference-in-Differences. Requires parallel trends.RDD — Regression Discontinuity at a running-variable cutoff.RKD — Regression Kink — change in slope, not level.Bunching — Estimator from mass concentrated near a threshold.PSM — Propensity Score Matching — matches treated/control on score.IPW/DR — Inverse-Probability Weighting / Doubly Robust.Event Study — Dynamic effects over time relative to an event.Synthetic Control — Builds a synthetic control from weighted units.LATE — Local ATE via IV — compliers only.Lee Bounds — Nonparametric bounds under sample selection.Mediation — Decomposes total effect into direct + indirect.
Missing / Bootstrap / Survey / Bayesian
Missing Data
What it is. Imputation and pattern analysis.
Options / Parameters
Simple — mean/median/mode — fast but understates SEs.KNN — Uses k nearest complete rows.MICE — Multiple imputation via chained equations — pooled with Rubin's rules.EM — Assumes multivariate normality.PMM — Predictive Mean Matching — preserves observed distribution.Patterns — Little's MCAR test + pattern visualization.Sensitivity — How results change under MNAR assumptions.
Bootstrap
What it is. Resampling for SE, CI, validation.
Options / Parameters
SE / CI — Bootstrap standard error / CI for a statistic.CI types — percentile / basic / BCa / studentized.Regression — Bootstrap coefficient estimates.Jackknife — Leave-one-out.Permutation (2-sample) — Exact test for difference of means.Bayesian bootstrap — Dirichlet weights — Bayesian variant.
Survey / Complex Samples
What it is. Design-based analysis (strata, clusters, weights).
Options / Parameters
Design — Specify strata / cluster / weights vars.Means / Totals / Proportions — Design-based estimates.DEFF — Design effect — variance inflation relative to SRS.Rao-Scott χ² — Chi-square corrected for complex design.Replicate Weights — BRR / JKn via replicate weight columns.
Bayesian Models
What it is. MCMC-based Bayesian inference (PyMC).
Options / Parameters
Linear / Logistic / Hierarchical — Base models.Draws / Tuning / Chains — MCMC settings (defaults 2000/1000/4).Posterior Predictive Check — Simulate data from posterior and compare to observed.Prior vs Posterior — Visualize how beliefs updated.Credible Interval — 94% HDI per parameter.Bayes Factor — Compare evidence between two models.MCMC Diagnostics — R-hat, ESS, trace plots.
ML & SHAP
ML Models
What it is. Random Forest, XGBoost, SVM, Decision Tree, Gradient Boosting.
Options / Parameters
Hyperparameters — Per-model — n_estimators, max_depth, etc.Train/Test Predict — Fit on train, predict on held-out test (test_size configurable).Tree Diagram Export — Export decision tree as Graphviz DOT for visualization.
ML Evaluation
What it is. Performance metrics.
Options / Parameters
Cross-Validation — K-fold CV (default 5).ROC / AUC — Binary classification performance.Confusion Matrix — TP/FP/TN/FN counts.Feature Importance — From tree-based models.
SHAP
What it is. Explainability: per-feature contribution per prediction.
Works with xgboost / random_forest / decision_tree.
Niche
Niche (DSGE / Conjoint / Neural / MIDAS)
What it is. DSGE macro, Conjoint analysis, Deep NN, MIDAS (mixed-frequency data).
Accessible via Command Palette (Ctrl+K) for the less common ones.