Problem
Putting language models into safety-critical operations raises questions that accuracy benchmarks do not answer. If a safety filter says a prompt is safe, will it still say so when the prompt is changed slightly? When a model is adapted to a new task, what has the added component learned to remember, and can that be bounded?
Both are questions about dynamics. State space models (SSMs), the recurrent architecture behind Mamba and S4, are dynamical systems, so the tools used to certify controllers apply to them directly: reachability, contraction and model reduction. This project uses those tools to build language-model components whose behavior can be checked, not just measured.
Certifiable safety classifiers
A safety head is a small classifier that screens prompts before a model responds. The question is whether it can promise not to change its mind.
- Can a safety classifier prove its decision holds under small input changes? For a linear SSM classifier, the full set of scores it could produce, while every token embedding is nudged within a fixed bound, can be computed exactly. That set stays bounded, whatever the prompt length, if and only if the state transition is contracting. A simple training penalty enforces contraction, and on toxic-comment data it raises the share of certified decisions from 41% to 59%, with a sharp transition exactly where the theory predicts (AIMS @ COLM 2026).
- Does the idea carry over to jailbreak detection? A 54,000-parameter, contraction-regularized S4 head reads only the prompt’s token embeddings from Mamba-130M, so it runs in about 5 ms on a CPU, before any response is generated. Trained on JailbreakBench, it detects all six attack families and transfers zero-shot to AdvBench and HarmBench.
- What does the guarantee cover, and what not? A linear probe on the same embeddings detects jailbreaks just as well, so the head’s value is the certificate, which no probe provides. The certificate holds in embedding space: adaptive attackers who swap whole tokens can still evade the head, which sets the next problem.
Adapters with analyzable memory
Parameter-efficient fine-tuning such as LoRA corrects each token’s representation independently, so the adapter itself has no memory. Tasks such as long-document question answering need an evolving summary of what has been read.

LoRA (left) adds a low-rank correction to a frozen weight matrix. The Hankel reduced-order model (HRM) adapter (right) adds a small state space model beside the frozen MLP, whose state carries information across tokens. From SSM Adapters via Hankel Reduced-order Modeling, HiLD @ ICML 2026.
- Can a fine-tuning adapter carry state? The HRM adapter is a small SSM added beside the MLP of a frozen transformer. Because its dynamics are linear and time-invariant, the recurrence runs as an FFT convolution, with wall-clock cost on par with LoRA (HiLD @ ICML 2026).
- What should a small memory keep? Directions of the state that the input can excite and that also affect the output. These are ranked by Hankel singular values from Gramians estimated on the data, and balanced truncation keeps only the important ones, with a certified bound on the error this introduces. The spectrum doubles as a fingerprint of how much memory a task needs.

Left: on a finite-state-machine tracking task, HRM and its balanced-truncation reduction (HRM-BT) beat LoRA at every sequence length. Right: the Hankel singular value spectrum shows that only 5 of 32 state directions matter for this task. From SSM Adapters via Hankel Reduced-order Modeling, HiLD @ ICML 2026.
- Where in the network should memory go? The injection site decides which tasks an adapter suits: adapters in attention help retrieval, while an adapter beside the MLP integrates information over the sequence.
- Does it help at scale? On Mistral-7B, at the same 8.4 million trainable parameters as LoRA, AdaLoRA, DoRA and QLoRA, HRM gives the best results on the tasks that need sequential integration (QuALITY and QMSum), and trails on retrieval-heavy NarrativeQA.
Outcome
- Safety filters that can prove their decisions. A training recipe that makes a lightweight safety classifier certifiably stable under bounded input perturbations, with the threshold predicted by theory.
- Jailbreaks caught before the model answers. Zero-shot, the head flags over 98% of held-out harmful prompts from AdvBench and HarmBench, from the prompt alone.
- Fine-tuning that remembers. A stateful adapter that beats LoRA-family methods on long-document tasks, including 35% higher relative accuracy on QuALITY with Mistral-7B at an equal parameter budget.
- Published and open-sourced. Papers at the HiLD workshop at ICML 2026 and the AIMS workshop at COLM 2026, with code for both.
