Recent Research by CAIAC Members
LLM CoTs Remain Monitorable When Being Unfaithful Requires Computation
Arav Dhoot, Yixiong Hao
July 15, 2026
Sample-Level White-Box Detection of Alignment Faking
Lakshya Chaudhry, Tianqin Meng, Anthony Nguyen, Yashraj Panwar, Yuqi Sun, Zhuofan Ying
July 10, 2026
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
Pranav Mehta
July 3, 2026
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
Niklas Weller, Emilio Barkett
May 24, 2026
Representation Without Control: Testing the Realization Effect in Language Models
Ciarán Walsh, Emilio Barkett
May 24, 2026
Sparse Autoencoder Interpretability of the METAGENE-1 Genomic Foundation Model
Mannat Vikramaditya Jain, Peyton Jackson, Bridget Liu, Ciaran Walsh, Astrid Teo
April 27, 2026
The AI in the Mirror: LLM Self-Recognition in an Iterated Public Goods Game
Olivia Long, Carter Teplica
August 25, 2025
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
Emilio Barkett, Olivia Long, Paul Kröger
August 3, 2025
Efficiently Detecting Hidden Reasoning with a Small Predictor Model
Rohan Subramani, Vishnu Vardhan Sai Lanka, Yau-Meng Wong, Daria Ivanova
July 13, 2025
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
Emilio Barkett, Olivia Long, Madhavendra Thakur
June 12, 2025
The Partially Observable Off-Switch Game
Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, Scott Emmons
April 11, 2025
SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models
Carter Teplica, Yixin Liu, Arman Cohan, Tim GJ Rudner
December 15, 2024
Adaptive Contextual Perception: How to Generalize to New Backgrounds and Ambiguous Objects
Zhuofan Ying, Peter Hase, Mohit Bansal
December 2, 2024
Generalization Analogies (Genies): A Testbed for Generalizing AI Oversight to Hard-to-Measure Domains
Joshua Clymer, Garrett Baker, Rohan Subramani, Sam Wang
November 13, 2023