You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Principled Agents is a research nonprofit working on defining goals that are safe for powerful AI to pursue. We are looking to cover operational expenses for up to one year, starting from December 2026.
The Principled Agents agenda is to formally specify safety properties that collectively make a utility function safe to optimize. The target we have in mind is unmanipulated and fully informed human approval, pursued corrigibly and conservatively.
We take a full-stack approach, starting with conceptual desiderata, then mathematically defining them alongside the utility functions and setups that incentivize them, and finally testing them empirically as a reward signal. Individual projects we run tend to focus on a single safety property (e.g. corrigibility), and involve reviewing existing proposals, identifying where they fail, and proposing new mechanisms that address at least some of the issues.
This funding would be used to cover operational expenses at our current level, allowing us to continue making progress in our research.
Our monthly expenses are just under $40,000. Approximately 70% of this is allocated to the salaries of three researchers. The remaining 30% covers compute for experiments (10%), office space (8%), and operational overhead along with miscellaneous expenses (12%).
Rubi Hudson is the founder of Principled Agents and the research lead. He has a track record of producing research that makes progress on important theoretical problems in AI safety. Notable outputs include Joint Scoring Rules (AAAI 2025), which showed how to elicit predictions from agents without incentivizing them to affect outcomes, and Corrigibility Transformation (ICML 2026), which introduced a method for modifying goals to remove instrumental incentives for avoiding updates and shutdown. Rubi is wrapping up a PhD at the University of Toronto, previously participated in the MATS program, and has mentored for the SPAR and PRISM programs.
Baram Sosis and Artem Petrov are researchers at Principled Agents. Baram has a PhD in Mathematics from the University of Pittsburgh, and previously participated in PIBBSS and the Anthropic Fellows program, where he investigated the effects of helpful-only training on emergent misalignment. Artem has three years of experience as a software engineer, including two years at Palisade Research where he led a team evaluating AI cyber capabilities.
Principled Agents has a number of projects to pursue that we believe could lead to a positive impact on making AI safer. It is likely that at least some individual projects fail due to misguided approaches or because the underlying problem is particularly intractable. We aim to determine this early on, by focusing on building formal models to capture the important intuitions or reveal inconsistencies.
The broader project can fail overall if there exist crucial safety properties that are very resistant to formalization. If neither we, nor other groups pursuing similar research, can successfully characterize a target, then aligning superintelligence will rely on ad hoc measures. However, partial progress can still reduce the number of issues that need to be addressed without a principled solution.
The organization can fail if employees leave to pursue other opportunities, or if we run out of funding. The impact of this would depend on whether current employees are able to pursue similar research in their subsequent work.
Principled Agents started with a grant of £132,000 from UK AISI through The Alignment Project, and has since received an additional $20,000 from the Corrigibility Research Fund. Together, these grants cover our expenses until the end of November 2026.