Yves
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Fine-tuning LLM agents to cooperate in multi-agent settings, building on my Cooperative AI Research Fellowship (top ~1% of 1,100+ applicants), targeting an ICLR 2027 submission in September 2026.
AI systems are increasingly deployed as interacting agents that negotiate, trade, and share resources. This creates safety risks that single-agent alignment does not address miscoordination, conflict, and collusion between agents, as laid out in the Cooperative AI Foundation report "Multi-Agent Risks from Advanced AI" (Hammond et al. 2025) https://arxiv.org/pdf/2502.14143.
This project targets that gap: cooperative fine-tuning of LLM agents with reinforcement learning on game-theoretic social dilemmas (Prisoner's Dilemma, Stag Hunt, public-goods games), using rewards derived from moral principles, so that cooperative behavior generalizes to more complex multi-agent settings the model was never trained on. The headline benchmark is N-player resource management in GovSim (Piatti et al.) https://arxiv.org/pdf/2404.16698, where current LLM agents reliably collapse into short-sighted, self-interested behavior.
Further evaluation benchmarks CoopEval (cooperation-sustaining mechanisms) and GT-HarmBench (game-theoretic safety risk), so results are comparable against the current state of the field.
The closest prior work Moral Alignment in LLMs (Tennant et al. 2025) showed that intrinsic moral rewards can teach a small LLM to behave morally in the Iterated Prisoner's Dilemma, but only against a single fixed opponent with fixed payoffs. My goal is cooperation that actually transfers to complex N-player games. I'm attacking it on three fronts:
Reward design: solving the reward challange of finding a game-agnostic reward that drives cooperation across different games simultaneoiusly.
Environment and curriculum: finding the mix of games, opponents (fixed strategies, learning opponents, self-play) and training paradigme, that makes cooperation generalize. Including generalization from 2-player matrix games to N-player common games.
Algorithm: solving credit assignment in iterated play, where vanilla GRPO dilutes one advantage over a whole multi-turn trajectory. I'm currently testing per-step and return-to-go advantage variants and a very promising self-distillation approach (SDPO).
~$8k living stipend for 2 months Aug 2026 – Sept 2026. ($4k/month, modest for Zurich where I'm based). Final experiments, writing, and the ICLR 2027 submission (Sept 25).
+ $4k living stipend for 1 month Okt 2026. For open-sourcing the environments, training pipeline, models and follow up experiments for comunity and collaborators.
+ $8 living stipend for 2 months Nov 2026 – Dec 2026. Rebuttal experiments and extending the results for the camera-ready version if accepted. Strengthened resubmission if not accepted initially.
I do this work unpaid at ~90% of my time; my savings ran out in July and I can cover my bills until the end of August 2026. This grant lets me finish what I started instead of abandoning it for a full-time job search and side-hustles.
I work on this solo, supervised by Prof. Zhijing Jin and her research group (ETH Zürich / Vector Institute / University of Toronto). She was my mentor during the Cooperative AI Research Fellowship. After the fellowship ended, I kept working on the project without any funding.
I hold an MSc in Physics from ETH Zurich (2025) with a machine learning focus. During my studies I produced three publications: two NeurIPS 2024 workshop papers (RL and optimization for robotics control; knowledge distillation for symmetry invariance in neural networks) and a journal paper in Physics in Medicine & Biology (multi-objective optimization in radiation oncology). The project itself is off to a strong start: I built the full post-training stack myself, the RL fine-tuning pipeline (GRPO, SDPO), the game environments, and the evaluation harness.
The real risk is generalization to complex N-player games. Transfer hasn't worked yet with the current single-turn, 2-player setup, which is not suprising. The harder pieces (multi-turn credit assignment, richer curricula, the full capability of GRPO/SDPO reasoning fine-tuning) aren't trained and fully exploited yet. That's precisely what this grant funds.
If generalization doesn't hold once those are in place, the fallback is a systematic transfer study plus open-source environments and pipelines the community can build on. The main downside is missing the ICLR deadline, in which case the work moves to ICML/NeurIPS/ACL.
9000$ as the stipend from my Cooperative AI Research Fellowship. Since the fellowship ended, nothing: applications to the Long-Term Future Fund, Coefficient Giving, and BlueDot were unsuccessful this year.
There are no bids on this project.