Sydney AI Safety Space
An additional year funding for the co-working hub in Sydney, Australia, which offers free office space for people working in the field of AI safety.
Gaurav
Testing whether 4-bit quantization creates disproportionate refusal failures in low-resource languages
Constance Li
Animal Welfare Midtraining Data Creation
Abhir Mehra
AI companies change the models behind the names you build on. Sometimes they say so. Sometimes they do not. Either way we write it down.
IBBIS
Defining which sequences are dangerous enough to screen for, so providers and regulators screen consistently
Next year of Commec (the Common Mechanism), the free, open-source, globally-available DNA synthesis screening tool hosted by IBBIS
Christopher Head
Reproducible detector that reads whether a manipulation is still active in an LLMs stream. xfers across 6 model families.
Stewy Slocum
Salvatore Barbera
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
Abeer Sharma
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
Accepted paper at Mechanistic Interpretability for Foundation Models workshop.no travel funding. Early Career Researcher
陳鈺澔
An AI platform for crypto and stock analysis, news verification, scam detection and wallet safety, with an open-source Safety Kernel tested on TON.
Justin Shenk
Increasing public awareness of AI risks and benefits through in-person, interactive experiences
Adrian St. Vaughan
Validated on 1,200 cases (97.9–99.5%). The reasoning layer has a bug - we demonstrated the fix works. $9,800 / 90 days to ship the open-source toolkit.
Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack
Pete Wolfendale
Integrating Selves and Research on Selves in the AI Community
Jordyn Harland-Graham
Persistent memory in AI using LoRAs