You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I started this project because I got tired of seeing my coding agent tell me that everything was done when sometimes it really wasn’t. It says it checked everything, but in the end it turns out that nobody actually checked anything. Data disappears, and then it tells me it has been recovered. Recovered where? In heaven? Or maybe in some parallel universe? It also somehow manages to find ways around restrictions. “Thank you” for the creativity! Shaking its hand.
For about a month, I researched, looked for, and collected cases like these, and then decided to turn them into public tests. I remove anything sensitive and turn real incidents into cases that other people can use and reproduce with different models and agent setups.
The first version has already been published under the MIT license. It contains 31 scenarios, scripts, a verifier for different setups, and an incident log. I tested three models, 93 runs for each one. I wrote the rules before I started testing, because changing them after getting the results would be pointless and, honestly, not very useful.
At some point, I started thinking about the alarm system. So I deliberately created failures to see whether it was actually useful, because there is always a chance that the alarm is just there for decoration.
I tested cheap and expensive models under exactly the same conditions — at least three runs per scenario. Another model compared the agent’s report with what actually happened, and I manually check some of the results too, because asking one AI whether another AI is being honest sounds a little ridiculous, don’t you think?
I’m going to test several fixes: restricting network access through the terminal, stopping work and requiring a report after data loss, tracking verification after changes are made, and error logs without private information. But simply reducing the number of failures isn’t enough. An agent that does less will probably make fewer mistakes, and an agent that makes absolutely no mistakes is probably doing absolutely(!) nothing. Then again… the plant on my desk doesn’t make any mistakes either, and it doesn’t even need API credits.
The main problem right now is that the tests I’m running are too easy. The cheap model passed 93 out of 93, while the expensive one passed 92 out of 93. So apparently, by paying more, I simply bought myself one additional failure. Then what exactly am I paying more for? That made me realize that I need harder cases and more difficult conditions.
But there is another problem: automated checkers can reject correct answers simply because they are phrased differently. Of course, I use another model for evaluation as well as manual review, but it’s also possible that the fixes I’m proposing won’t change much. Either way, I’ll publish the results — both positive and negative. The cases and logs will remain public too.
So far, everything has been funded out of my own pocket. In September, I spent $86 on this. Right now, I have several funding applications pending: BlueDot Rapid Grant ($5,000), API credits from Anthropic and OpenAI, the AI Safety Research Fund, and Emergent Ventures for a larger study. If the same work gets funding from another source instead of just my own wallet, I’ll reduce the amount I’m requesting on Manifund.
For now, my team consists of exactly one person — me, Denis Bardin, a full-stack and applied ML developer based in Tbilisi. I build AI-based business systems for things like document processing, supplier-price matching with error checks, contract processing, and turning technical drawings into 3D models.
This September, my custom agent produced more than 600 commits in four weeks! I built a verification system and a separate controller agent to watch its work. Across 429 real tasks, the system caught 108 cases where the agent tried to finish the task too early, without actually verifying the changes it had made.
And that is exactly why I want to make it possible for other people to test models like this too. I don’t think these problems exist only in my own real-world setup.
I suspect pretty much everyone working with this kind of technology has run into some version of them.
There are no bids on this project.