You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Benchmarks have become the de-facto legible surface for AI models, but there is no accountability for labs or eval shops. They self-attest numbers, trust-me-bro their procedures, and hope that nobody notices that they were lying. I want to change that.
I've been auditing model benchmarks for their reproducibility and validity. As of this summer, I've been producing verifiable receipts to the most popular coding benchmarks voluntarily without pay. That work has just been validated by OpenAI. In February, they recommended SWE-bench Pro as the standard coding capability bench, and everyone followed. That's when they were topping the charts. In July, with Fable in the lead, OpenAI deprecated it in public. They produced headline numbers that cannot be reconstructed from anything they published. Their methodology is full of vibes. In contrast, my audit proved that at least 15.0% of tasks are defective. Every label in the audit is verified and reproducible by anybody with model access. Theirs you have to trust, mine you can run.
This legibility has been the backbone of my research program as an independent, and I believe this accountability should be widely visible so that they can do their own science.
SWE-bench Pro was one of eight benchmarks I audited. Between the audits, critical flaws were found where keys that fail their own graders, measure the wrong thing, or ignore harmful side effects such as rm -rf'ing the user contents. Recently, AVERI audit institute priced their cheapest audit program at $300,000 per engagement, which I demonstrated can be done by just me.
I intend to keep running the cheap audits, and make it standard practice. Each audit will:
1. Rick a bench
2. Run the checks against its own shipped artifacts
3. Publish every finding with a re-runnable receipt
4. Give the authors right of reply, and;
5. File the actionable part where the maintainers work
Each audit produces an implemented upstream fix to its grader (e.g. harbor-framework/harbor#2266). Receipts do not remove judgment. They expose its inputs, so another auditor can reproduce or contest it. Audits that come back clean get published too.
Then compound it: the checklist becomes a disclosure standard for what an eval must publish, the checks become tools anyone can run (my determinacy tool already audits benchmarks it wasn't written for), and the failure modes get filed into the standing registries evaluators already read (first filing: BenchRisk/BenchRisk#8, the NeurIPS 2025 benchmark-reliability registry). The full argument is in Assurance at the Boundary (DOI 10.5281/zenodo.21461148), my response to the AVERI agenda.
Tranched, so one funder can make the first move and the rest can co-fund. $10k: the next 10 or so audits, targets chosen by citation weight, receipts and right of reply included; covers run costs and two months of full-time work. $30k: 30 audits total, plus the disclosure standard v1 (record schemas for what an eval must publish), the checklist hardened into tools a second team runs without me, and one documented outside reproduction of a sampled audit verdict. $75k: 12 months full time: 100 audits committed, the standard, each with a ten-task discovery-benchmark pilot to replace what the Pro deprecation vacated (post-cutoff, contamination-controlled by construction; twenty tasks is stretch, and so is extending the same replay standard to AI agents' claims about their own work). Every tranche ends in a public, reproducible artifact. Breakdown at $75k: 82% stipend (taxes included), 8% compute and run costs, 10% buffer.
Solo. I work full-stack: the artifacts (audit repos, tools, upstream fixes, DOIs), the methodology (the checklist, right of reply, preregistration), the argument (Assurance at the Boundary), and the epistemics behind it (Verifiable Knowledge, the declared standard the audits grade against). The track record is the eight audits and where they landed: an upstream grader fix implemented, a failure mode filed into a Zenodo registry, an external convergence when OpenAI deprecated the benchmark I had receipted a month earlier. Everything is at june.kim and github.com/kimjune01, reproducible from committed inputs. I design for a distrusting auditor by default.
Audits are demand-constrained. The risk is that nobody with stakes reads them. SWE-bench Verified's contamination was common knowledge while it stayed the field's reported number for two years. Mitigations are built into the shape of the work (findings filed as fixes where maintainers work, failure modes filed into registries evaluators read). I record the outcome and move to the next target. A null result is a possible outcome of the grant, and reporting it is part of the work.
$0 in the last 12 months. The eight audits, the papers, and the tooling were all unfunded. I submitted an application to the Long-Term Future Fund today (20 July 2026) for the same program ($75k, pending); I will update this page when it resolves.