Research at Lennox Digital
Our research methodology is anchored in mechanistic understandability, empirical verification, and provable safety bounds. We operate across four dedicated research programs.
Mechanistic Interpretability
Neural models compute by propagating activations through high-dimensional attention and MLP blocks. We develop automated tools to extract discrete, human-interpretable sub-circuits responsible for reasoning, factual recall, and goal formulation.
Training massive dictionary models to disentangle polysemantic neurons into monosemantic concepts.
Applying steering vectors directly during inference to inhibit deceptive or harmful reasoning pathways.
Automating the extraction of minimal computation subgraphs for specific tasks and safety bounds.
Autonomous Agent Safety Bounds
When models are granted execution tools, shell access, and external API permissions, the failure modes transition from text hallucinations to real-world state corruption. We engineer dynamic invariants that mathematically bound agent behavior.
Cryptographically verifiable runtime sandboxes that monitor tool calls against formal policy models.
Defining safety envelopes that terminate agent execution if predicted system divergence exceeds threshold limits.
Studying emergent dynamics and collusive failure modes across interconnected autonomous swarms.
Scalable Oversight & Red Teaming
As models surpass human ability in specialized domains, human evaluators can no longer spot subtle flaws, deceptive hallucinations, or strategic compliance. We research AI-assisted auditing protocols.
Using verifiable, constrained models to audit step-by-step reasoning chains of frontier architectures.
Automated adversarial agents generating evolutionary attacks targeting covert biases and failure modes.
Distinguishing genuine alignment from strategic sycophancy during reinforcement learning.
Lennox Verification Engine (LVE)
Our flagship open-source software initiative. LVE is a high-performance evaluation library designed for parallel activation extraction, automated adversarial probing, and verifiable model scoring.
Read our empirical findings
Our research papers and technical reports are available openly to researchers and institutions worldwide.
