You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Research laboratories are giving more and more tasks to AI agents in their field of research. They will need some form of monitoring to ensure safety: an LLM "judge" looking over what the agent is doing, or a probe that will look at its activations... to catch if there's something dangerous going on. What if the agent tries to argue?
In my previous work I have shown that a simple (and slightly worrying) law applies: social pressure displaces the judgment of a model by a constant fraction towards what the pressure wants it to. The constants I fitted on one experiment predicted a different experiment without any refitting (R² 0.73/0.69 out of sample). Currently under review at ICLR 2027.
So can we argue our way around monitors? Can the agent trick the judge by using flattery, act as if they have power over the judge, create a false authorization, impersonate the operator or insert malicious commands into its response? Are there cases where the monitor will allow certain things to pass by? That's what I want to test in this project.
I'll try these 5 attacks on LLM judges and activation probes designed to catch malicious behavior. I'll then test some simple countermeasures such as rephrasing the agent's output using another model we trust, combining a judge and probe and fine tuning our monitor to be resistant to persuasion.
Paper C: corrupting monitors (Month 1-4). By how much do LLM judges and probes give in to these 5 types of pressure? How much does it cost to move a probe versus a judge? How can we fix the behaviour of the system? And a public acceptance test for monitors. As mentioned previously I have a tool for this: on Llama-3.1-8B social pressure will move the judgement of a model by a constant fraction towards what it is pressured to think. Constants can be transferred to a new campaign. Go/no-go decision in Month 2: if previous work already covers this law for monitors, this paper will focus on the cost of the attacks and on the test.
Paper A, diffuse sabotage (month 1-7). Using a simulation of a sabotage in a research setting (Terekhov et al, arXiv 2606.08892) with 8B model organisms, I will compare black box and white box monitoring per action and over time. Go/no-go decision will be made in Month 4: if white box over time cannot beat the best black box benchmark on recall at a 1% false positive rate (95% CI), then I will release a negative result. Deliverables: Two papers on arXiv. With code, data, model organisms, probes and acceptance tests open source. First part already released at github.com/lciric/controle-ia
Three goals to be achieved by the end of the twelve months:
The minimum, $11,000, pays for two months full-time. $22,000 pays for four months, enough to finish the monitor paper (Paper C), which stands on its own. The full goal, $58,630, pays for seven months (December 2026 to June 2027): $35,000 for my stipend ($5,000 a month before French income tax and social contributions), $4,000 of GPU compute, $7,000 of AI research tooling, $4,000 for paid human checks of the ground truth, $2,500 for one workshop trip, $800 of miscellaneous costs and a $5,330 contingency.
Over those seven months there are two things I would like to accomplish, and a third afterwards if more funding comes in. Write a paper about corrupting monitors (Paper C). How far can monitors be persuaded by an adversary? Which solutions will stand? This will include the red team findings, different solutions compared and all code necessary to reproduce the experiments. This should be finished by month 4-5.
Write a paper about detecting diffuse sabotage (Paper A) using my over time monitoring technique. I hope to show that it can detect patterns that "per action" monitors cannot. And that it can stand against an adversary that knows how it's being detected.
If everything goes well and more funding comes in, I would like to expand this to multiple agents afterwards (Paper B). Can we take an agent and split its changes up across multiple instances of the same model and fool the monitor? I set go/no-go criteria before running anything. If the monitors hold up against these attacks, that is also a great result and I will release that as well. I'm not trying to reach a certain conclusion...
Just me. I have a background in physics (ENS Paris-Saclay), computational neuroscience (ENS; Master's thesis in Srdjan Ostojic's group. Published in PLOS Comp Bio in 2023) and Machine Learning (Mines Paris-PSL). Did a research internship at IBM research in Zurich. Publication in Neuromorphic Computing and Engineering in 2022.
Since March 2026 I've been researching AI safety on my own dime. Here are some of my findings:
Discovered a displacement law for sycophancy on Llama-3.1-8B. Pressure will displace the score given by a model by a constant fraction towards the desired value. R² 0.77/0.79 in sample. Can predict new campaigns with no further readjustment with R² 0.73/0.69. This paper is under review at ICLR 2027.
I also found that what a behavior tracking probe detects makes up <1% of the causality that affects the model's answer, whereas the full hidden state makes up 68%. Meaning probes can find behaviour without having influence over it. This is very important when trying to supervise using interpretability techniques.
I found a mistake in my own work and flagged it. The model's responses were being stored cut off at 300 characters, which led me to believe that I had found a safety effect when it was just a generic shift in distribution. I wrote about my findings and published the correction.
My current project controle-ia has 98 papers that have been hashed and saved to the source directory, and 3,497 auto-verified quotes with 0 failures. I have an activation-logging harness validated on Llama-3.1-8B and 602 automated tests. I believe I have used around $20 worth of compute so far.
I'm in the final stage of selection for the 2026 Anthropic Fellows program in AI safety. I'm French, from Paris.
Most likely my attacks will not be good enough. If I attack the model by trying to convince it in simple ways and the monitor can detect that then I haven't done much. A real world agent could find many more ways to do this. I have a few attacks that are more subtle (inserting messages into the prompt and spoofing system messages) that try to do more than just socially pressure the model, and I will be honest about the scope of the test. The proceed/abort criteria I set in advance help with this: if the monitor does not get corrupted, I will publish that result without making the test sound harder than it was.
Risk #2: Open model results may not apply to newer models that I do not have access to inspect. If the API credits I've applied for come through, I can use similar "LLM-judge" type attacks on closed API models. However to probe these models I would need their weights. This is a legitimate concern that I will disclose.
Risk #3: I may not be able to complete both papers in seven months. I will publish my monitor corruption paper first as it's further along. I can always release my diffuse sabotage paper as a work in progress/preprint if I run out of time.
One last risk is personal: this project is self-funded.
None towards this project in the past 12 months. I've currently spent about 7500€ of my own money on this since March 2026. I currently have applications pending with EA Funds (Transformative AI Fund), BlueDot (2 tracks), Foresight, Lightcone, GTR, Middle Powers, the CLR Fund, Emergent Ventures, the AI Alignment Foundation and Iliad. Also applications for compute and API credits with Anthropic, OpenAI, NDIF and Google's TPU Research Cloud. None have made a decision yet. I will only take one stipend for any given period.