neighbayes.models.SDEMPanelRE

class neighbayes.models.SDEMPanelRE(**kwargs)[source]

Bayesian spatial Durbin error panel model with unit random effects.

\[y_{it} = X_{it}\beta + (WX)_{it}\theta + \alpha_i + u_{it}, \quad u_{it} = \lambda (Wu)_{it} + \varepsilon_{it}\]

Combines the SDEM mean structure (covariates plus their spatial lags) with random unit effects \(\alpha_i \sim N(0, \sigma_\alpha^2)\) and a spatially-correlated error term governed by \(\lambda\).

Parameters:
formula : str, optional

Wilkinson-style formula, e.g. "y ~ x1 + x2". Requires data, unit_col, and time_col.

data : pandas.DataFrame, optional

Long-format panel data when using formula mode.

y : array-like, optional

Stacked response of shape (N*T,). Required in matrix mode.

X : array-like or pandas.DataFrame, optional

Stacked design matrix. Required in matrix mode.

W : libpysal.graph.Graph or scipy.sparse matrix

Spatial weights of shape (N, N). Used to construct the WX block and the spatial filter on the disturbance. Should be row-standardized.

unit_col : str, optional

Column in data identifying the cross-sectional unit. Required in formula mode.

time_col : str, optional

Column in data identifying the time period. Required in formula mode.

N : int, optional

Number of cross-sectional units. Required in matrix mode.

T : int, optional

Number of time periods. Required in matrix mode.

priors : dict, optional

Override default priors. Supported keys:

  • lam_lower (float, default -1.0): Lower bound of Uniform prior on \(\lambda\).

  • lam_upper (float, default 1.0): Upper bound of Uniform prior on \(\lambda\).

  • beta_mu (float, default 0.0): Normal prior mean for \([\beta, \theta]\).

  • beta_sigma (float, default 1e6): Normal prior std for \([\beta, \theta]\).

  • sigma2_alpha (float, default 2.0): InverseGamma prior alpha for sigma2

  • sigma2_beta (float, default var(y)): InverseGamma prior beta for sigma2 for \(\sigma\).

  • sigma_alpha_sigma (float, default 10.0): HalfNormal prior std for \(\sigma_\alpha\).

  • nu (float, default 4.0): Fixed Student-t degrees of freedom (only used when robust=True).

logdet_method : str, optional

How to compute \(\log|I - \lambda W|\); auto-selected when None (default).

robust : bool, default False

If True, replace the Normal innovation with Student-t.

w_vars : list of str, optional

Names of X columns to spatially lag. By default all non-constant columns are lagged. At least one column must be lagged; if no WX columns remain a ValueError is raised. Pass a subset to restrict which variables receive a spatial lag.

Notes

The base-class model argument is not exposed; pooled mean structure (model=0) is used because unit heterogeneity is captured by the random effect rather than by within-unit demeaning.

__init__(**kwargs)[source]

Methods

__init__(**kwargs)

fit([draws, tune, chains, random_seed, ...])

Draw samples from the posterior for the panel model.

fitted_values()

Return fitted values at posterior mean parameters.

residuals()

Return residuals y - fitted_values.

spatial_diagnostics()

Run Bayesian LM specification tests and return a summary table.

spatial_diagnostics_decision([alpha, format])

Return a model-selection decision from Bayesian LM test results.

spatial_effects([return_posterior_samples])

Compute Bayesian inference for direct, indirect, and total impacts.

summary([var_names])

Return posterior summary table.

Attributes

inference_data

Return the ArviZ InferenceData from the most recent fit.

pymc_model

Return the PyMC model object built for the most recent fit.

fit(draws=2000, tune=1000, chains=4, random_seed=None, progressbar=True, sampler=None, gibbs_backend='auto', thin=1, n_jobs=-1, idata_kwargs=None, **sample_kwargs)[source]

Draw samples from the posterior for the panel model.

Mirrors SpatialModel.fit(): dispatches to this model’s Gibbs sampler (sampler="gibbs") or NUTS (sampler="nuts"). When sampler is None (default), Gibbs is used if the model has a registered Gibbs sampler (Gaussian FE families), otherwise NUTS. (The two fit bodies are duplicated across the cross-section and panel base classes until Phase 5c collapses the hierarchies.)

Parameters:
draws : int

Post-warmup draws, warmup steps, and number of chains.

tune : int

Post-warmup draws, warmup steps, and number of chains.

chains : int

Post-warmup draws, warmup steps, and number of chains.

random_seed : int, optional

Seed for reproducibility.

progressbar : bool, default True

Show progress bar(s) during sampling.

sampler : {"gibbs", "nuts", None}, default None

Sampling method. None auto-selects Gibbs when this model has one, else NUTS.

gibbs_backend : {"auto", "jax", "numpy"}, default "auto"

Execution backend for the Gibbs sampler. "auto" uses JAX when installed and supported, otherwise NumPy. Ignored for NUTS.

thin : int, default 1

Keep every thin-th post-warmup Gibbs draw (Gibbs only).

n_jobs : int, default -1

Parallel workers for the NumPy Gibbs path (Gibbs only).

idata_kwargs : dict, optional

{"log_likelihood": True} stores the complete Jacobian-corrected pointwise log-likelihood that az.loo / az.waic / az.compare need, for Gibbs and NUTS alike. Off by default, as in PyMC: it holds one value per draw, chain, and observation (16 GB at n = 250,000 with 4 × 2,000 draws). For NUTS the dict is also passed to pm.sample.

**sample_kwargs

For NUTS, forwarded to pm.sample (target_accept, nuts_sampler="blackjax"/"numpyro"/"nutpie", …). For Gibbs, the family’s declared options (slice_width, …); an unsupported key raises.

Returns:

Posterior samples and diagnostics.

Return type:

arviz.InferenceData

fitted_values()[source]

Return fitted values at posterior mean parameters.

Returns:

Posterior-mean fitted values (on the model’s native scale; fixed-effects-transformed for panel models).

Return type:

np.ndarray

property inference_data : arviz.data.inference_data.InferenceData | None[source]

Return the ArviZ InferenceData from the most recent fit.

Returns:

The inference data object, or None if the model has not been fit yet.

Return type:

arviz.InferenceData or None

property pymc_model : pymc.model.core.Model | None[source]

Return the PyMC model object built for the most recent fit.

Returns:

The model object used by fit(), or None if the instance has not been fit yet.

Return type:

pymc.Model or None

residuals()[source]

Return residuals y - fitted_values.

Returns:

Residual vector y - fitted_values on the same scale as fitted_values().

Return type:

np.ndarray

spatial_diagnostics()[source]

Run Bayesian LM specification tests and return a summary table.

Looks up the diagnostic suite registered for this model class and calls each test function on this fitted model, collecting the results into a tidy DataFrame. The set of tests depends on the model type — for example, an OLS model runs LM-Lag, LM-Error, LM-SDM-Joint, and LM-SLX-Error-Joint, while an SAR model runs LM-Error, LM-WX, and Robust-LM-WX. Panel models run the Panel--prefixed analogues (e.g. Panel-LM-Lag).

Requires the model to have been fit (.fit() called) and a spatial weights matrix W to have been supplied at construction time.

Returns:

DataFrame indexed by test name with columns:

Column

Description

statistic

Posterior mean of the LM statistic

median

Posterior median of the LM statistic

df

Degrees of freedom for the \(\chi^2\) reference

p_value

Bayesian p-value: 1 - chi2.cdf(mean, df)

ci_lower

Lower bound of 95% credible interval (2.5%)

ci_upper

Upper bound of 95% credible interval (97.5%)

The DataFrame has attrs["model_type"] (class name) and attrs["n_draws"] (total posterior draws) metadata.

Return type:

pandas.DataFrame

Raises:
  • RuntimeError – If the model has not been fit yet.

  • ValueError – If no spatial weights matrix W was supplied.

See also

spatial_diagnostics_decision

Model-selection decision based on the test results.

spatial_effects

Posterior inference for direct/indirect/total impacts.

Examples

>>> ols = OLS(formula="price ~ income + crime", data=df, W=w)
>>> ols.fit()
>>> ols.spatial_diagnostics()
                 statistic  median  df  p_value  ci_lower  ci_upper
LM-Lag                3.21    2.98   1    0.073      0.12      8.54
LM-Error              5.67    5.34   1    0.017      0.34     12.10
LM-SDM-Joint          7.89    7.12   4    0.096      1.23     18.32
LM-SLX-Error-Joint    6.45    5.98   4    0.168      0.89     15.67
spatial_diagnostics_decision(alpha=0.05, format='graphviz')[source]

Return a model-selection decision from Bayesian LM test results.

Implements the decision tree from Koley and Bera [2024] (the Bayesian analogue of the classical stge_kb procedure in Anselin et al. [1996]). Panel models use the Panel--prefixed test analogues and the panel decision specs, following Elhorst [2014]. The decision logic depends on the current model type and the pattern of significant tests:

From OLS (6-test decision tree):

  1. If only LM-Lag is significant → SAR.

  2. If only LM-Error is significant → SEM.

  3. If both are significant → use the Anselin–Florax / Koley–Bera robust pair: Robust-LM-Lag → SAR, Robust-LM-Error → SEM, both → SARAR. If neither robust test is significant, fall back to the lower raw p-value.

  4. If neither naive test is significant → OLS.

From SAR (3-test decision tree):

  • LM-Error significant → SARAR; LM-WX significant → SDM; Robust-LM-WX significant → SDM.

From SEM (2-test decision tree):

  • LM-Lag significant → SARAR; LM-WX significant → SDEM.

From SLX (4-test decision tree):

  • Robust-LM-Lag-SDM significant → SDM; Robust-LM-Error-SDEM significant → SDEM; both → MANSAR; neither → SLX.

From SDM: LM-Error-SDM significant → MANSAR; else SDM.

From SDEM: LM-Lag-SDEM significant → MANSAR; else SDEM.

Parameters:
alpha : float, default 0.05

Significance level for the Bayesian p-values.

format : {"graphviz", "ascii", "model"}, default "graphviz"

Output format. "model" returns the recommended-model name string. "ascii" returns an indented box-drawing rendering of the full decision tree with the chosen path highlighted. "graphviz" returns a graphviz.Digraph object that renders inline in Jupyter; if the optional graphviz package is not installed a UserWarning is issued and the ASCII rendering is returned instead.

Returns:

Recommended model name when format="model", an ASCII tree string when format="ascii", or a graphviz.Digraph when format="graphviz" (with ASCII fallback on missing dep).

Return type:

str or graphviz.Digraph

See also

spatial_diagnostics

Compute the Bayesian LM test statistics.

References

Koley and Bera [2024], Anselin et al. [1996], Elhorst [2014]

spatial_effects(return_posterior_samples=False)[source]

Compute Bayesian inference for direct, indirect, and total impacts.

Computes impact measures for each posterior draw, then summarizes the posterior distribution with means, 95% credible intervals, and Bayesian p-values. This is the fully Bayesian analog of the simulation-based approach in LeSage and Pace [2009] and the asymptotic variance formulas in Arbia et al. [2020].

Models without a spatial lag on y do not exhibit global feedback propagation through \((I-\\rho W)^{-1}\). However, models with spatially lagged covariates (SLX, SDEM) can still have non-zero neighbor spillovers captured in the indirect term.

Parameters:
return_posterior_samples : bool, optional

If True, return a (DataFrame, dict) tuple where the dict contains the full posterior draws under keys "direct", "indirect", and "total". Default False.

Returns:

If return_posterior_samples is False (default), returns a DataFrame indexed by feature names with columns for posterior means, credible-interval bounds, and Bayesian p-values.

If return_posterior_samples is True, returns (DataFrame, dict) where the dict has keys "direct", "indirect", "total", each mapping to a (G, k) array of posterior draws.

Return type:

pd.DataFrame or tuple of (pd.DataFrame, dict)

summary(var_names=None, **kwargs)[source]

Return posterior summary table.

Parameters:
var_names : list, optional

Variable names to include in the summary.

**kwargs

Additional arguments passed to arviz.summary().

Returns:

Posterior summary statistics.

Return type:

pandas.DataFrame