Anthology · Issue 04 · May 2026
We build capable, safe,
and interpretable AI.
Vellum is an applied AI safety lab. Our work focuses on alignment, constitutional methods, and the long-horizon behavior of frontier systems. The papers below are written for working researchers — and for the public that pays for the field.
Recent papers
All publications →May 4, 2026Methods
Constitutional methods at scale: a five-year retrospective
Ortiz · Kim · Avraham · Park · Reyes
Apr 18, 2026Evaluation
On the difficulty of evaluating long-horizon agent behavior
Park · Reyes · Caspi
Mar 30, 2026Interpretability
Mechanistic interpretability of multi-step planning
Caspi · Yoon · Aldred · Singh
Mar 11, 2026Alignment
Specification-gaming under reward modeling pressure
Aldred · Singh
Feb 22, 2026Safety
Calibrated refusal: an empirical study
Reyes · Avraham