27 September 2026 · LessWrong

Why I expect AI replication incidents by 2027

I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely.

12 February 2025 · with Jacob Haimes · Apart Research

Hunting for AI Hackers: LLM Agent Honeypot

Our current results indicate that LLM hacking agents exist but are in the very early stages of technology adoption for mass hacking.

7 January 2025 · LessWrong

Predicting AI Releases Through Side Channels

I recently explored whether we could predict major AI releases by analyzing the Twitter activity of OpenAI's red team members.

5 March 2024 · Substack

Replication Work of “How Johnny Can Persuade LLMs to Jailbreak Them”

Analysis of study about effectiveness of human persuasion to LLM jailbreaking.