duckLM user's guide
July 14, 2026 · View on GitHub
Task-oriented documentation for the duckLM regression macros. For a one-page signature reference see CHEATSHEET.md; for install/setup see the README.
- Fitting
- Predicting & prediction intervals
- Evaluating
- Inference: standard errors, p-values, confidence intervals
- Influence diagnostics
- Multiclass (multinomial)
- Tuning hyperparameters (cross-validation)
- Categorical features
- Contract & fine print
Fitting
{logit,linreg,poisson,gamma,tweedie,nbinom}_fit fit a regression of outcome
on every other column of the input table.
CREATE TABLE churn_model AS
SELECT * FROM logit_fit('training_data', 'churned');
CREATE TABLE rev_model AS
SELECT * FROM linreg_fit('sales', 'revenue');
CREATE TABLE claims_model AS
SELECT * FROM poisson_fit('policies', 'n_claims');
CREATE TABLE severity_model AS
SELECT * FROM gamma_fit('claims', 'claim_amount');
-- Tweedie for zero-inflated positive data (insurance pure premium):
-- power=1.5 is the compound Poisson-Gamma; p=1 is Poisson, p=2 is Gamma
CREATE TABLE pure_premium AS
SELECT * FROM tweedie_fit('policies', 'loss_cost', power := 1.5);
-- negative binomial for overdispersed counts (variance > mean);
-- alpha is the fixed dispersion (variance = mu + alpha*mu^2), alpha->0 = Poisson
CREATE TABLE visits_model AS
SELECT * FROM nbinom_fit('patients', 'n_visits', alpha := 0.5);
-- optional ridge regularization on any family
CREATE TABLE churn_model_reg AS
SELECT * FROM logit_fit('training_data', 'churned', l2 := 0.1);
SELECT * FROM rev_model;
-- ┌─────────────┬─────────────┐
-- │ feature │ coefficient │
-- │ (Intercept) │ 3.4097 │
-- │ ad_spend │ -2.0046 │
-- │ headcount │ 0.7201 │
-- └─────────────┴─────────────┘
| argument | default | meaning |
|---|---|---|
tbl | — | table/view name as a string (resolved via query_table; schema-qualified names work) |
outcome | — | column to predict, as a string; for logit_fit it must be 0/1 or boolean, for poisson_fit non-negative, for gamma_fit strictly positive |
max_iter | 50000 | hard cap on solver iterations (only the gd path gets near it; IRLS converges in ~5–10) |
learning_rate | NULL | step size on the standardized scale; NULL auto-picks a convergent default (4/(d+1+4·l2) logistic, 1/(d+1+l2) otherwise; Poisson/Gamma steps are additionally damped each iteration by the largest curvature weight, since theirs is unbounded) |
tol | 1e-10 | stop early when the gradient step is smaller than this |
l2 | 0.0 | ridge penalty (l2/2)·Σβ² added to the mean loss of the internally standardized problem, intercept unpenalized |
l1 | 0.0 | lasso penalty l1·Σ|β| (feature selection); combine with l2 for elastic net. Intercept unpenalized |
offset_col | NULL | name of a column holding a per-row offset added to the linear predictor η = offset + xβ with a fixed coefficient of 1 (not fit, not penalized) |
weights_col | NULL | name of a column of non-negative per-row sample weights — the loss (and internal standardization) are weighted by them |
solver | 'auto' | 'auto' = IRLS, falling back to gradient descent on a rank-deficient design (see below); 'gd' = force Nesterov gradient descent; 'irls' = force Fisher-scoring IRLS (supports l2, l1, offset, weights; needs full-rank features) |
Coefficients are on the original feature scale (features — and for
linear/Poisson/Gamma regression the outcome — are rescaled internally only
for optimizer conditioning), so unpenalized results match R's glm()/lm(),
statsmodels, or unpenalized scikit-learn up to convergence tolerance.
Solver. The single-outcome families default to solver := 'auto', which
picks the fastest solver that is valid for the problem:
- IRLS (Fisher scoring / iteratively reweighted least squares). It solves a
weighted least-squares system each iteration and reaches the exact maximum-
likelihood estimate in ~5–10 iterations rather than the ~1000 gradient descent
needs — 3–10× faster end to end for the same answer, matching statsmodels to
machine precision. Without
l1the system is solved by a pure-SQL matrix inverse; withl1there is no inverse to take, so the same normal equations go to cyclic coordinate descent with soft-thresholding instead (this is glmnet's proximal-Newton scheme), which also makes lasso coefficients exactly zero rather than merely small. - Nesterov gradient descent as an automatic fallback when
XᵀWXturns out to be singular — a constant or perfectly collinear feature, or complete separation in logistic regression. GD degrades gracefully on those inputs where IRLS cannot.
The fallback is free: each solver is a recursive CTE gated so that the one not in
play emits no seed row and never iterates. So 'auto' costs the same as IRLS on
a well-conditioned fit, and the same as GD when it has to fall back.
Force a solver with 'gd' or 'irls'. Forcing 'irls' on a singular design
raises a clear error rather than falling back — useful if you would rather hear
about a rank-deficient design than have it silently absorbed.
-- the default already uses IRLS here; this is just explicit
SELECT * FROM poisson_fit('policies', 'n_claims', solver := 'irls');
Offset / exposure. An offset is a known per-row term in the linear
predictor — most often log(exposure) for a Poisson/Gamma rate model (claims
per policy-year, events per person-time). Pass the column name via
offset_col, and pass the same offset_col to the matching *_predict /
*_evaluate so scoring includes it. Matches R's offset= /
statsmodels' GLM offset=.
-- claims modelled per unit of exposure: E[claims] = exposure · exp(xβ)
CREATE TABLE m AS
SELECT * FROM poisson_fit('policies', 'n_claims', offset_col := 'log_exposure');
SELECT * FROM poisson_predict('m', 'new_policies', offset_col := 'log_exposure');
Negative binomial. nbinom_fit(tbl, outcome, alpha := 1.0, ...) models
overdispersed counts (variance = μ + α·μ², so variance > mean). alpha is the
fixed dispersion; α→0 recovers Poisson. Matches statsmodels
GLM(..., family=NegativeBinomial(alpha=α)). Composes with offset, weights,
and ridge/lasso. To estimate α, nbinom_dispersion(tbl, outcome, alpha_grid)
returns the profile log-likelihood per α — the argmax is the (grid-resolution)
MLE dispersion; feed it back into nbinom_fit:
SELECT alpha FROM nbinom_dispersion('patients', 'n_visits', [0.1, 0.25, 0.5, 1.0, 2.0])
ORDER BY loglik DESC LIMIT 1; -- the estimated dispersion
For a sharper estimate without a huge grid, nbinom_dispersion_refine re-sweeps
a fine grid around the peak automatically (two-stage refinement).
Sample weights. weights_col names a column of non-negative per-row
weights; the loss and the internal standardization are weighted by them.
Matches scikit-learn's sample_weight and R's weights=. Integer weights
behave exactly like replicating each row that many times. Weights apply to
fitting only — *_predict and *_evaluate don't take them.
CREATE TABLE m AS
SELECT * FROM linreg_fit('survey', 'income', weights_col := 'sampling_weight');
Tweedie. tweedie_fit(tbl, outcome, power := 1.5, ...) takes a variance
power p that unifies the log-link families: p=1 is Poisson, p=2 is Gamma,
and 1<p<2 is the compound Poisson-Gamma that models data with exact
zeros and positive continuous values together — the classic insurance
pure-premium / loss-cost use case. tweedie_predict returns exp(score);
tweedie_evaluate(model, tbl, outcome, power := 1.5) returns Tweedie deviance,
deviance-based pseudo-R², and Pearson dispersion. Outcomes must be ≥ 0 for
1≤p<2 (strictly positive for p≥2). Matches
TweedieRegressor(power=p, alpha=0, link='log').
Ridge semantics. Like glmnet, the penalty applies to standardized
coefficients, so a given l2 has comparable strength regardless of feature
or outcome scale, and the intercept is never penalized. Exact scikit-learn
equivalents (fit on z-scored features, outcome transformed as the macro does
internally): Ridge(alpha = n*l2) for linear, LogisticRegression(C = 1/(n*l2)) for logistic, PoissonRegressor(alpha = l2) / GammaRegressor( alpha = l2) on the mean-scaled outcome. A small l2 also gives perfectly
separable logistic data a finite, fast solution.
Lasso & elastic net. l1 adds an L1 penalty that drives coefficients to
exactly zero for feature selection (FISTA / proximal-gradient); combine
l1 and l2 for elastic net. Available on every family (the intercept is
never penalized). Linear matches Lasso(alpha = l1) /
ElasticNet(alpha = l1+l2, l1_ratio = l1/(l1+l2)) on the standardized problem.
-- keep only the features that matter
CREATE TABLE sparse_model AS
SELECT * FROM linreg_fit('wide_table', 'target', l1 := 0.1);
SELECT feature FROM sparse_model WHERE coefficient != 0; -- selected features
Predicting
{logit,linreg,poisson,gamma,tweedie,nbinom}_predict score a table with a
fitted model, matching model features to columns by name. Extra columns
(ids, the outcome itself, …) are passed through untouched and ignored by the
scoring.
-- logistic: adds prob DOUBLE and pred BOOLEAN (pred = prob >= threshold)
SELECT * FROM logit_predict('churn_model', 'new_customers');
SELECT * FROM logit_predict('churn_model', 'new_customers', threshold := 0.7);
-- linear: adds prediction DOUBLE
SELECT * FROM linreg_predict('rev_model', 'pipeline');
-- poisson: adds prediction DOUBLE = exp(score), the expected count
SELECT * FROM poisson_predict('claims_model', 'new_policies');
-- gamma: adds prediction DOUBLE = exp(score), the expected value
SELECT * FROM gamma_predict('severity_model', 'open_claims');
-- tweedie: adds prediction DOUBLE = exp(score), the expected value
SELECT * FROM tweedie_predict('pure_premium', 'renewals');
-- negative binomial: adds prediction DOUBLE = exp(score), the expected count
SELECT * FROM nbinom_predict('visits_model', 'new_patients');
Prediction intervals (*_predict_ci) add a confidence band on the
predicted mean to the point prediction. The coefficient covariance
φ·(XᵀWX)⁻¹ is estimated from the training table, giving each row's linear
predictor a standard error SE(η̂) = √(x̂ᵀ Cov x̂); the band is g⁻¹(η̂ ± crit·SE)
back on the response scale (prediction, conf_low, conf_high). It matches
R's predict(..., se.fit=TRUE) / statsmodels get_prediction(). Pass the model,
the training data (for the covariance) and its outcome, and optionally a
newdata table to score (default: the training data); conf_level, offset_col
and weights_col behave as in the fit:
CREATE TABLE model AS SELECT * FROM poisson_fit('policies', 'n_claims');
-- fitted means with a 95% band, in-sample:
SELECT * FROM poisson_predict_ci('model', 'policies', 'n_claims');
-- forecast new rows with a 99% band:
SELECT * FROM poisson_predict_ci('model', 'policies', 'n_claims', newdata := 'new_policies', conf_level := 0.99);
The interval is on the mean (like the coefficient CIs, it is z-based for
fixed-dispersion families and Student-t for the estimated-dispersion ones). A
singular XᵀWX yields a finite prediction but NULL band.
Evaluating
{logit,linreg,poisson,gamma,tweedie,nbinom}_evaluate score tbl with a fitted
model and its outcome column and return a one-row table of goodness-of-fit
metrics. Pass the training table for in-sample fit or a holdout for
out-of-sample. Metrics follow the standard statsmodels / scikit-learn
definitions (verified against both).
SELECT * FROM linreg_evaluate('rev_model', 'sales', 'revenue');
-- ┌───────┬────────┬────────┬────────┬─────────┬──────────┬─────────┬─────────┐
-- │ n │ rmse │ mae │ r2 │ adj_r2 │ loglik │ aic │ bic │
-- └───────┴────────┴────────┴────────┴─────────┴──────────┴─────────┴─────────┘
SELECT * FROM logit_evaluate('churn_model', 'training_data', 'churned');
-- n, accuracy, auc, log_loss, loglik, deviance, null_deviance, pseudo_r2, aic, bic
| macro | metrics returned |
|---|---|
linreg_evaluate | n, rmse, mae, r2, adj_r2, loglik, aic, bic |
logit_evaluate | n, accuracy, auc, log_loss, loglik, deviance, null_deviance, pseudo_r2, aic, bic |
poisson_evaluate | n, rmse, mae, loglik, deviance, null_deviance, pseudo_r2, aic, bic |
gamma_evaluate | n, rmse, mae, deviance, null_deviance, pseudo_r2, dispersion |
tweedie_evaluate | n, rmse, mae, deviance, null_deviance, pseudo_r2, dispersion |
nbinom_evaluate | n, rmse, mae, loglik, deviance, null_deviance, pseudo_r2, dispersion, aic, bic |
pseudo_r2 is McFadden's (logistic) or deviance-based (Poisson/Gamma); aic
and bic use k = number of model coefficients (intercept included). Gamma's
log-likelihood/AIC depend on the dispersion parameter, so it reports deviance,
deviance-based pseudo-R², and the Pearson dispersion instead.
Inference
{logit,linreg,poisson,gamma,tweedie,nbinom,multinom}_summary return a
coefficient table with standard errors, test statistics, p-values and
confidence intervals — the equivalent of R's summary(glm(...)) or statsmodels
.summary(), in pure SQL. Pass a fitted model and its training data (same shape
as *_evaluate):
CREATE TABLE model AS SELECT * FROM poisson_fit('visits', 'n');
SELECT * FROM poisson_summary('model', 'visits', 'n');
| feature | coefficient | std_error | statistic | p_value | conf_low | conf_high |
|---|---|---|---|---|---|---|
| (Intercept) | 0.5190 | 0.0429 | 12.09 | 3e-33 | 0.4348 | 0.6031 |
| x1 | 0.7843 | 0.0338 | 23.17 | 0 | 0.7180 | 0.8507 |
| … |
The covariance matrix is Cov(β̂) = φ · (XᵀWX)⁻¹, with W the per-family
expected-information (Fisher) IRLS weights evaluated at the fitted coefficients
— so the standard errors are identical to statsmodels to machine precision.
It composes with offset_col and weights_col (analytic/var_weights
convention), and conf_level sets the interval (default 0.95):
SELECT feature, coefficient, conf_low, conf_high
FROM gamma_summary('model', 'claims', 'cost', weights_col := 'exposure', conf_level := 0.99);
Robust & cluster-robust standard errors. Pass robust := 'hc0' (or 'hc1',
'hc2', 'hc3') for heteroskedasticity-consistent sandwich SEs
(XᵀWX)⁻¹ B (XᵀWX)⁻¹ — valid when the variance model is misspecified
(over/under-dispersion, heteroskedasticity). Pass cluster_col for one-way
cluster-robust SEs (correlated observations within a group — panel/repeated
measures). Both are dispersion-free and z-based, use the observed-information
bread (so they match statsmodels cov_type='HC0'/'cluster' to machine
precision for every family), and compose with weights and offsets:
-- heteroskedasticity-robust
SELECT * FROM poisson_summary('model', 'claims', 'n', robust := 'hc0');
-- clustered by region (correlated within region)
SELECT * FROM poisson_summary('model', 'claims', 'n', cluster_col := 'region');
hc0 is the White estimator; hc1 applies the n/(n−d) correction; hc2/hc3
divide the squared residual by (1−hᵢ) / (1−hᵢ)² using the GLM leverage
hᵢ (hc3 is the recommended default for small samples). Cluster-robust uses
the Stata finite-sample factor (G/(G−1))·((n−1)/(n−d)).
The reference distribution follows R's glm/lm convention: the statistic
is a z-score with a normal p-value/CI when the dispersion is fixed
(logit/poisson/nbinom), and a t-score with n − d degrees of freedom
when the dispersion is estimated (linreg/gamma/tweedie, Pearson φ). Both
the normal and Student-t CDFs/quantiles are computed in pure SQL and exposed as
reusable helpers — norm_cdf(z), norm_ppf(p), t_cdf(t, df), t_ppf(p, df)
— matching SciPy to ~1e-12 across the practical range (t_ppf is exact for
df = 1; extreme-tail quantiles at very low df are approximate):
SELECT norm_ppf(0.975) AS z95, t_ppf(0.975, 30) AS t95; -- 1.959964, 2.042272
Notes and limits: p-values/CIs assume an unpenalized (maximum-likelihood)
fit — they are not valid for l1/l2-penalized models (the SE ignores the
penalty). Collinear, constant, or severely ill-conditioned features make XᵀWX
singular/indefinite, so every std_error/statistic/p_value/CI is returned
as NULL rather than a fabricated value (the coefficient column is still
reported); the same happens for the estimated-dispersion families when residual
df n − d ≤ 0. Degenerate inputs return NULL, never an error.
Multinomial inference is multinom_summary(model, tbl, outcome, conf_level := 0.95).
It inverts the baseline-category softmax Fisher information (the block matrix
I[(c,j),(c',k)] = Σᵢ pᵢ_c(δ_cc' − pᵢ_c')·xᵢⱼxᵢₖ, dispersion fixed at 1, so
z inference like logistic) and returns one row per estimated (class, feature)
for the K − 1 non-reference classes — matching statsmodels MNLogit to machine
precision (and, for two classes, exactly reproducing logit_summary). The
reference class (alphabetical-minimum label) is the fixed baseline and is not
reported.
CREATE TABLE model AS SELECT * FROM multinom_fit('iris', 'species');
SELECT * FROM multinom_summary('model', 'iris', 'species');
Influence diagnostics
{logit,linreg,poisson,gamma,tweedie,nbinom}_influence(model, tbl, outcome)
return, for every training row, the five standard case-deletion diagnostics
alongside the input columns — for spotting outliers and high-leverage points:
| column | meaning |
|---|---|
hat | leverage hᵢ = aᵢ·hwᵢ·xᵢᵀ(Xᵀdiag(a·hw)X)⁻¹xᵢ (GLM hat-matrix diagonal, in [0,1], summing to d) |
pearson_resid | Pearson residual (y−μ)/√V(μ) |
deviance_resid | signed-root deviance residual sign(y−μ)·√(unit deviance) |
std_resid | studentized (standardized Pearson) residual r_P/√(φ(1−h)) |
cooks_distance | Cook's distance (r_P²/φ)·h / (d(1−h)²) |
These match statsmodels GLMInfluence to machine precision (leverage and
Cook's distance use the observed-information hat matrix, as statsmodels does).
Rows near cooks_distance > 4/n or with high hat are the influential ones.
SELECT * FROM poisson_influence('claims_model', 'policies', 'n_claims')
ORDER BY cooks_distance DESC LIMIT 10; -- the 10 most influential observations
Multiclass
multinom_fit / multinom_predict / multinom_evaluate implement multinomial
(softmax) logistic regression for a categorical outcome with any number of
classes. It uses the identifiable baseline-category parameterization — one
coefficient set per class relative to a reference (the alphabetical-minimum
label, held at 0) — matching R's nnet::multinom and statsmodels MNLogit.
CREATE TABLE species_model AS
SELECT * FROM multinom_fit('iris', 'species');
-- (class VARCHAR, feature VARCHAR, coefficient DOUBLE); the reference class
-- has all-zero coefficients, so scoring is a self-contained softmax.
SELECT * FROM multinom_predict('species_model', 'new_flowers');
-- adds pred VARCHAR (argmax class) and probs MAP(VARCHAR, DOUBLE);
-- get a class probability with probs['setosa']
SELECT * FROM multinom_evaluate('species_model', 'iris', 'species');
-- (n, accuracy, log_loss)
The outcome is the class-label column (any type); every other column is a
numeric/boolean feature — dummy-encode categoricals first (below). multinom_fit
takes optional l2 / l1 (ridge / lasso / elastic-net, penalizing each class's
standardized coefficients, intercepts excluded); offset and weights aren't
available for multinomial. Coefficient standard errors are available via
multinom_summary.
Tuning hyperparameters
cv_l2 / cv_l1 / cv_power / cv_alpha run k-fold cross-validation over
a grid, returning one row per grid value with the mean held-out deviance
(squared error for linear). Pick the smallest cv_deviance.
All k × |grid| models are fit together in a single pass over the data, and
— as with *_fit — by IRLS, so a sweep costs a handful of iterations rather than
the thousands gradient descent needs. That includes cv_l1: the penalised models
are solved by coordinate descent inside the same iteration, and a grid mixing
l1 = 0 with penalised values simply solves each model by whichever method fits.
A rank-deficient design (constant or perfectly collinear feature) sends the whole
run back to gradient descent automatically, exactly as solver := 'auto' does for
a single fit.
| macro | tunes | families |
|---|---|---|
cv_l2(tbl, outcome, family, l2_grid, k := 5) | ridge l2 | linear/logistic/poisson/gamma |
cv_l1(tbl, outcome, family, l1_grid, k := 5) | lasso l1 | linear/logistic/poisson/gamma |
cv_power(tbl, outcome, power_grid, k := 5) | Tweedie power | (tweedie) |
cv_alpha(tbl, outcome, alpha_grid, k := 5) | neg-binom alpha | (nbinom) |
SELECT * FROM cv_l2('training_data', 'churned', 'logistic', [0.0, 0.01, 0.1, 1.0])
ORDER BY cv_deviance LIMIT 1; -- the best l2
SELECT * FROM cv_power('claims', 'loss_cost', [1.2, 1.4, 1.6, 1.8]) ORDER BY cv_deviance LIMIT 1;
It's genuinely pure SQL: all k × |grid| models are fit simultaneously in one
recursive CTE (each fold-model's gradient sums only over its non-held-out
rows, with its own hyperparameter), then the held-out rows are scored.
Standardization is global (matching cv.glmnet), and folds are assigned
deterministically as (row# − 1) % k — shuffle the table first if its rows are
ordered by the outcome. Cost scales with k · |grid| · features · rows · iterations, so keep the grid modest.
Two-stage refinement
Rather than pay for one dense grid, sweep a coarse grid and then a fine
one zoomed in on the winner. reg_grid(lo, hi, n) builds an evenly spaced grid
(log_spaced := true for a geometric one), and each cv_*_refine /
nbinom_dispersion_refine wrapper runs the whole two-stage sweep in a single
call: it fits the coarse grid, finds the best value, then re-sweeps n_refine
(default 10) points bracketing that value between its two coarse-grid
neighbours, returning the refined curve. Take its argmin cv_deviance (or
argmax loglik) as the estimate. Two n-point stages resolve the optimum about
as finely as one n²-point grid at a fraction of the cost.
| macro | coarse → refined |
|---|---|
cv_l2_refine / cv_l1_refine(tbl, outcome, family, grid, k := 5, n_refine := 10) | ridge/lasso l2/l1 |
cv_power_refine(tbl, outcome, grid, k := 5, n_refine := 10) | Tweedie power |
cv_alpha_refine(tbl, outcome, grid, k := 5, n_refine := 10) | neg-binom alpha |
nbinom_dispersion_refine(tbl, outcome, grid, n_refine := 10) | dispersion (profile likelihood) |
-- coarse log grid, then auto-refine around the best ridge penalty
SELECT * FROM cv_l2_refine('training_data', 'churned', 'logistic', reg_grid(1e-3, 10, 8, log_spaced := true))
ORDER BY cv_deviance LIMIT 1;
SELECT * FROM nbinom_dispersion_refine('patients', 'n_visits', [0.1, 0.5, 1.0, 2.0, 4.0])
ORDER BY loglik DESC LIMIT 1; -- sharper dispersion estimate than the coarse grid
Assumes the coarse grid brackets the optimum; if the best lands on an endpoint the refined grid is one-sided toward the interior — widen the coarse grid if you expect the optimum beyond it.
Categorical features
The fit macros treat every column as numeric (booleans become 1/0). To use a
categorical/VARCHAR column you dummy-encode it first — and
dummy_encode_sql(tbl, outcome) writes that SQL for you. It one-hot encodes
every VARCHAR column except the outcome with R-style treatment contrasts
(k−1 indicators per factor, dropping the first level as the reference); numeric
and boolean columns pass through untouched. The result reproduces
lm(y ~ ... + C(factor)) to ~1e-8.
Because a macro can't return a data-dependent set of columns, it returns the
SELECT as text — run it as a second step (trivial from any driver):
sql = con.sql("SELECT dummy_encode_sql('sales', 'revenue')").fetchone()[0]
con.sql(f"CREATE TABLE encoded AS {sql}")
con.sql("SELECT * FROM linreg_fit('encoded', 'revenue')")
-- e.g. dummy_encode_sql('sales', 'revenue') returns:
SELECT * EXCLUDE (region),
(region = 'North')::INT AS "region_North",
(region = 'South')::INT AS "region_South",
(region = 'West')::INT AS "region_West" -- 'East' is the reference
FROM sales
Interactions and transforms are still plain columns you add yourself
(ln(x) AS log_x, a * b AS a_x_b, …). A NULL category yields NULL dummies, so
that row is dropped by the fit (as R drops NA). tbl must be a table/view
(resolvable in duckdb_columns).
Contract & fine print
- Feature columns must be castable to
DOUBLE(numeric or boolean). Booleans become 1/0. - Training drops rows with a
NULLoutcome or anyNULLfeature. A feature column that is entirely NULL is rejected with a clear error (as R'sglm()and scikit-learn do) rather than silently excluded. - Prediction returns
NULLoutputs for rows where a model feature isNULLor the column is missing entirely. - Zero-variance (constant) feature columns get coefficient
0. A constant outcome inlinreg_fitis fine (intercept = mean, slopes 0). A constant column makesXᵀWXsingular, so the default'auto'solver falls back to gradient descent here. - Exactly collinear features (singular design) don't error under the default
'auto'solver: it detects the singularXᵀWX, falls back togd, and converges to one valid least-squares solution. Its predictions match any other solver's; the individual coefficient split across the collinear columns is arbitrary (as it is for every solver). Forcingsolver := 'irls', and*_summary/*_predict_ci/*_influence, instead return a clear error / NULL because the covariance is undefined there. - Data with no finite maximum-likelihood solution — perfectly separable
logistic data, or an all-zero-count Poisson/all-constant Gamma outcome —
still terminates with large coefficients rather than erroring, but only
after running all
max_iteriterations. (IRLS diverges to ~1e305 on separable data, which the'auto'solver treats as a failure and falls back togd, so you get the same large-but-finite coefficients as before.) For separable logistic data a positivel2gives a finite, fast solution (it penalizes the diverging slopes). For an all-zero Poisson outcome it's the intercept that diverges, andl2does not penalize the intercept (by design, matching glmnet/sklearn), so lowermax_iterinstead. - Guarded name collisions (clear errors, never silent misbehavior): column
and table names beginning with
__reg_are reserved everywhere, checked case-insensitively; a feature named(Intercept)is rejected at fit time (as the outcome it's fine);prob/predcolumns (any case) are rejected bylogit_predictandpredictionbylinreg_predict/poisson_predict/gamma_predict— drop them first withSELECT * EXCLUDE (...). - Prediction matches model features to columns by exact, case-sensitive name string — score with the same column spellings you trained with.
- The training set is materialized as an in-memory list during optimization — comfortable up to a few hundred thousand rows × dozens of features.
__reg_*macros (e.g.__reg_fit,__reg_score,__reg_eval,__reg_summary) are internal helpers; call the public macros instead.