Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack
Yves
Reinforcement learning for cooperative LLM agents to address multi-agent risks.
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.
Pete Wolfendale
Integrating Selves and Research on Selves in the AI Community
Pip Foweraker
A nightmarishly hard AI safety strategy game about holding p(Doom) down. You can't win; you can only buy time.
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
Jai Dhyani
Creating conditions for cooperative strategies to dominate adversarial ones among near-future AIs while we still can
Isaiah Tapia
Researching Public Compute for California
Nickola Horozov
David Franklin
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Gaurav Hadavale
Seeking travel and registration fee support to present my sole-authored mechanistic audit of medical AI at MICCAI 2026 MI4MedFM workshop at Strasbourg, France
lucas.irwin
A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implem
Jesse Charlie