crabbymetrics
  • Home
  • API
    • API Overview
    • Regression And GLMs
    • Survival / Event-Time
    • Causal Inference And Panels
    • MPE_CBPS
    • Hypothesis Testing And Utilities
    • Transforms
    • Estimation Interfaces
  • Internals
  • Regression
    • OLS
    • ABC OLS
    • Anytime-Valid Confidence Sequences
    • Ridge
    • Bagged Polynomial Regression
    • Fixed Effects OLS
    • ElasticNet
    • Logit
    • Multinomial Logit
    • Poisson
    • MLE Prediction Interface
    • Survival / Recurrent Events
    • GMM
    • MEstimator Poisson
  • Causal / Panels
    • Balancing Weights
    • Chronos LTV Balancing
    • Cressie-Read And Rényi Balancing
    • EPLM
    • Average Derivative
    • Double ML And AIPW
    • Richer Regression
    • TwoSLS
    • Synthetic Control
    • Synthetic DID
    • Augmented Balancing For Panel Data
    • Horizontal Panel Ridge
    • Matrix Completion
    • Interactive Fixed Effects
    • Staggered Panel Event Study
    • Joint Hypothesis Tests
    • Dynamic Treatment Effects
  • Transforms
    • PCA And Kernel Basis
    • Sparse Factor Rotations
  • Ablations
    • Variance Estimators
    • Semiparametric Estimator Comparisons
    • Two-Period Semiparametric DID
    • Bridging Finite And Superpopulation
    • Panel Estimator DGP Comparisons
    • Same Root Panel Case Studies
    • Randomized Sketching And Least Squares
    • Estimator Scaling And References
  • Optimization
    • Optimizers
    • GMM With Optimizers
  • Ding
    • Chapter Index
    • Foundations (1-4)
    • Design And Adjustment (5-8)
    • Finite And Superpopulation (9)
    • Observational Studies (11-13, 27)
    • Instrumental Variables (21, 23)

On this page

  • 1 Where it fits
  • 2 Counting-process partial likelihood
  • 3 Implementation walkthrough
  • 4 Inference limitation
  • 5 Performance and numerical behavior
  • 6 Python API
  • 7 Minimal example
  • 8 summary() contract

AndersenGill

Counting-process Cox model for recurrent events or split risk intervals

1 Where it fits

Group: Survival / event-time models

AndersenGill extends the Cox partial likelihood to counting-process style risk intervals \((s_i, t_i]\). It is the natural entry point for recurrent events, delayed entry, and time-split records with time-varying covariates.

The fitted coefficients still live on the proportional-hazards log-risk scale, so the prediction contract mirrors CoxPH: latent log hazard ratio plus relative risk.

2 Counting-process partial likelihood

For interval record \((s_i,t_i]\), the risk set at event time \(t_i\) is

\[ R(t_i)=\{r:s_r<t_i\leq t_r\}. \]

The class maximizes

\[ \ell_p(\beta) = \sum_{i:d_i=1} \left[ x_i'\beta -\log\left\{\sum_{r\in R(t_i)}\exp(x_r'\beta)\right\} \right]. \]

Tied stop-time events use the Breslow denominator. Optimization shares CoxPH’s Newton solver, \(10^{-8}\) solve ridge, step clipping, and backtracking. Risk moments use stable log rescaling, not index clipping. Convergence requires either \(\|\nabla\ell_p\|_\infty<\tau\) or an accepted partial-likelihood change below \(\tau(1+|\ell_p|)\).

3 Implementation walkthrough

AndersenGill calls the same native partial-likelihood kernel and Newton loop as CoxPH, with one additional entry-time predicate.

  1. The wrapper validates aligned arrays, finite covariates, binary events, and every interval condition \(0\leq s_i<t_i\). It has no subject identifier and therefore cannot validate interval ordering or overlap within a subject.
  2. Slopes start at zero and covariates are centered internally. For each distinct stop time with events, the inner scan includes record \(r\) precisely when \(s_r<t_i\leq t_r\), matching the open-left, closed-right interval.
  3. The scan accumulates dynamically rescaled mass and covariate moments, with one numerator contribution and shared denominator per tied event. Likelihood-only backtracking skips derivatives. Validated owned fitting runs detached from the GIL.
  4. The shared Newton step explicitly inverts a Hessian with \(10^{-8}\) subtracted from its diagonal, clips step coordinates to \([-1,1]\), and backtracks through at most 30 halvings. Maximum-score or relative-objective convergence is required; the iteration cap is failure.
  5. Final covariance is the inverse observed information across interval rows. The code has no cluster-score accumulation stage, so repeated records from one subject are treated as separate contributions to this model-based information.
  6. Prediction ignores interval endpoints because the fitted object is only a relative-risk model: it returns \(X\hat\beta\) or \(\exp(X\hat\beta)\). It does not estimate a baseline hazard or a subject-specific event process.

The source-level difference from CoxPH is deliberately small, which makes the counting-process risk-set rule clear. It also makes the API limitation clear: without subject IDs, the implementation cannot supply the robust sandwich that usually makes Andersen-Gill inference useful.

4 Inference limitation

The returned covariance is only the inverse observed partial-likelihood information. The API has no subject identifier and cannot aggregate score residuals across multiple intervals belonging to the same subject.

Warning

For recurrent-event Andersen-Gill analysis, the usual robust covariance clusters intervals by subject. This class does not implement that covariance. Its reported standard errors and \(p\)-values treat interval rows through the model-based information and should not be presented as robust recurrent-event inference.

The class also does not validate that a subject’s intervals are nonoverlapping or consistently ordered because subject membership is unavailable. Default prediction is relative risk \(\exp(x'\hat\beta)\); no baseline hazard or survival curve is estimated. Reaching max_iterations raises ValueError and leaves the estimator unfitted. A successful summary exposes converged, iterations, termination_reason, and objective, where objective=-\ell_p(\hat\beta).

5 Performance and numerical behavior

Each distinct event time still scans every interval, costing about \(O(Unp^2)\) per derivative evaluation for \(U\) distinct event times, plus an \(O(p^3)\) solve. A complete start/stop sweep remains future work. Stable risk sums fix the old clipping inconsistency, but new-data relative risks can still overflow. Missing subject-cluster covariance remains a material limitation for recurrent-event inference.

6 Python API

Constructor: cm.AndersenGill

Use fit(x, start, stop, event). predict_lin(x) returns the log hazard ratio and predict_relative_risk(x) returns the exponentiated risk multiplier. The default predict(x) is relative risk.

Code
print(inspect.signature(cm.AndersenGill))
()
Code
cls = cm.AndersenGill
display(HTML(html_table(["Public method"], public_methods(cls))))
Public method
AndersenGill()
fit(self, /, x, start, stop, event, max_iterations=50, tolerance=1e-08)
predict(self, /, x)
predict_lin(self, /, x)
predict_log_hazard_ratio(self, /, x)
predict_relative_risk(self, /, x)
summary(self, /)

7 Minimal example

rng=np.random.default_rng(34)
x=rng.normal(size=(180,1)); rate=0.05*np.exp(0.6*x[:,0]); t_event=rng.exponential(1.0/rate); c=rng.exponential(18,size=180); stop=np.minimum(t_event,c); event=(t_event<=c).astype(float)
start=np.zeros_like(stop)
start_long=np.concatenate([start, stop/2]); stop_long=np.concatenate([stop/2, stop]); x_long=np.vstack([x,x]); event_long=np.concatenate([np.zeros_like(event), event])
model=cm.AndersenGill(); model.fit(x_long,start_long,stop_long,event_long)
fit=model.summary(); print({key: fit[key] for key in ['converged', 'iterations', 'termination_reason', 'objective']})
print(model.predict_lin(x[:5]))
print(model.predict(x[:5]))
{'converged': True, 'iterations': 3, 'termination_reason': 'Relative objective tolerance reached', 'objective': 335.91455804071126}
[-0.03003575 -0.94620885  1.93646531  0.3624618   0.48414704]
[0.97041084 0.38821    6.93419737 1.43686233 1.62279025]

8 summary() contract

The table below is generated by fitting the live class in this repository and then inspecting summary(). Shapes are shown because most values are plain NumPy arrays or scalars.

Code
rng=np.random.default_rng(134); x=rng.normal(size=(90,1)); rate=0.05*np.exp(0.5*x[:,0]); te=rng.exponential(1.0/rate); c=rng.exponential(15,size=90); stop=np.minimum(te,c); event=(te<=c).astype(float)
start=np.zeros_like(stop); start_long=np.concatenate([start, stop/2]); stop_long=np.concatenate([stop/2, stop]); x_long=np.vstack([x,x]); event_long=np.concatenate([np.zeros_like(event), event])
model=cm.AndersenGill(); model.fit(x_long,start_long,stop_long,event_long)
summary = model.summary()
display(HTML(html_table(["summary() key", "shape"], summary_shape_rows(summary))))
summary() key shape
model ()
coef (1,)
hazard_ratio (1,)
se (1,)
z (1,)
p_value (1,)
vcov (1, 1)
log_likelihood ()
converged ()
iterations ()
termination_reason ()
objective ()

crabbymetrics 0.9.0

 
  • v0.9 migration

  • Reproduction