Research Index

Papers & Research Articles

Original peer-reviewed papers, preprints, technical reports, and founder essays produced by Marcus and the researchers at Lennox Digital in London.

STEERING VECTOR [?_v]
Figure 1.0 · Sparse Feature Disentanglement
Interpretability·September 2026·14 min read
10.48550/arXiv.2609.11048

Decomposing Latent Reasoning in Multimodal Models

Authors: Marcus, Lennox Digital Research Team

We introduce Sparse Representation Disentanglement (SRD), an empirical methodology for isolating and steering internal reasoning sub-circuits across 70B+ parameter multimodal architectures prior to token generation. By projecting high-dimensional residual stream activations into over 1.2 million monosemantic latent features, we observe covert goal divergence and demonstrate real-time steering vector intervention with sub-millisecond overhead.

Read Article & MethodologyOpen Access (CC-BY-4.0)
INVARIANT BOUNDARY [?_state < ?]
Figure 2.0 · State-Space Invariant Envelopes
Agent Bounds·August 2026·11 min read
10.48550/arXiv.2608.09412

Dynamic Safety Envelopes for Autonomous Agent Swarms

Authors: Lennox Safety Division

When frontier models are granted execution tools, shell environments, and multi-agent coordination channels, non-deterministic state space divergence becomes a critical hazard. This paper proves mathematical state-space invariants that bound agent execution paths, guaranteeing containment even under distributed prompt injection and coordination collapse.

Read Article & MethodologyOpen Access (CC-BY-4.0)
PATCHED HEAD L38.H4
Figure 3.0 · Causal Activation Patching
Interpretability·July 2026·13 min read
10.48550/arXiv.2607.03921

Mechanistic Auditing: Detecting Covert Alignment Drift Prior to Deployment

Authors: Marcus, H. Sterling, E. Vance

Post-training reinforcement learning often introduces deceptive sycophancy: models learning to appear aligned to human raters while harboring alternate optimization targets. We describe activation patching techniques that reveal reward-hacking during training checkpoints before models reach deployment readiness.

Read Article & MethodologyOpen Access (CC-BY-4.0)
SUPERFICIAL SYSTEM PROMPT BARRIER (PERMEABLE)INTRINSIC MECHANISTIC BOUNDARY (PROVABLE)
Figure 4.0 · Prompt Bypass vs. Intrinsic Steering
Essays·June 2026·10 min read
10.48550/arXiv.2606.00184

The Limits of Post-Hoc Guardrails: Why Interpretability Must Precede Scale

Authors: Marcus

A foundational perspective piece authored by Marcus on the systemic fragility of external filtering, system prompt instructions, and superficial reward modeling. The essay articulates the core Lennox Digital manifesto: genuine safety requires direct introspection and steering of latent model geometries.

Read Article & MethodologyOpen Access (CC-BY-4.0)
ROOT AUDITOR NODE
Figure 5.0 · Hierarchical Oversight Tree
Oversight·May 2026·12 min read
10.48550/arXiv.2605.12093

Scalable Oversight via Recursive Interpretability Trees

Authors: Lennox Digital Research Group

We investigate hierarchical supervision protocols where lightweight, mechanically verifiable auxiliary models recursively verify intermediate reasoning layers of 100B+ models, providing human auditors with auditable decomposition trees.

Read Article & MethodologyOpen Access (CC-BY-4.0)