You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
We're requesting $53,900 for 6 months (1 December 2026 to 31 May 2027) to test an assumption that rarely gets checked. Oversight schemes (monitoring, verification, debate, ensembles) assume the model doing the checking fails independently from the one it's checking.
Separate developers often train on overlapping data with similar methods, and a single developer's models overlap more still. When checker and checked share a blind spot, the check passes the very mistakes it exists to catch.
About 600 tasks with checkable answers, some of them agent tasks whose tool calls are logged, go through 8 or more models from 5 or more developers, and every claim and tool call gets an error label. Pairwise tables then show, for each checker, the share of another model's errors it lets through next to what independent failure predicts.
The costly piece already runs: a platform sending live traffic to more than 100 models, logging every call. Licenses: CC BY 4.0 for data, Apache 2.0 for code.
The point is to see how often an AI model misses a mistake made by another one, especially when they come from the same company. That way, if you're building a system where one AI is supposed to watch another, you can actually know how independent they are.
Here's how we'd do it. We start by gathering roughly 600 jobs that have right answers you can look up, like cited question answering, pulling facts from documents, or running multi step tasks in a sandbox where every action gets logged. Next, each job goes to no fewer than 8 different models from at least 5 different builders, and we lock the version for each model we test. Then we break every answer down into single statements or tool calls and mark each one right or wrong against the reference answer. We use a hand labeled set of about 1,000 of these to see how much error our own automatic grading adds. After that, for every pair of models, we calculate how frequently the checker model fails to spot an error from the other one, and see whether that's higher than it would be if the two models erred independently. We slice this by whether the models are from the same developer or not, by how much better one is than the other, and by what kind of task it is. And we test one idea directly: that two smaller checkers from unrelated builders will catch more of a model's mistakes than one larger checker from that model's own developer.
What this buys for safety. A lot of plans for controlling powerful AI depend on using other models as guards, often weaker models watching stronger ones, or on groups of models. But if those guard models have the same blind spots as the one they're watching, then the safety you think you have is weaker than it looks, and you won't see the gap until it matters. By publishing these correlation numbers, we give researchers what they need to pick models for independence, not just raw performance, and it lets evaluators report both. No matter what we find, it's useful. If the errors aren't correlated, then AI checking holds up better than people fear. If they are strongly linked, then we know current setups need more variety to be safe.
How you'll know we did it. We release everything: the tasks, every model's output, all the labels for claims and tool calls, and all the code. The pairwise results get published with confidence intervals for each task category. We give a clear yes or no on the hypothesis. And we write a short report that says what didn't work just as plainly as what did.
Possible snags. Matching up individual claims across different answers is messy, which is why the plan includes the human labeled set to measure that noise. Model providers push updates all the time, which can change these relationships; we handle that by locking versions and dating every test run. It's also key to remember that these results only apply to tasks where answers can be fact checked, and we will be very clear about that limit.
We're asking for $53,900 to cover 6 months of work, from 1 December 2026 to 31 May 2027.
Here's what that pays for:
My time as founder and project lead, half time for the full 6 months: $25,000 (46%).
API calls for hosted models, at least 8 different models across roughly 600 tasks and agent runs: $9,000 (17%).
People to label the calibration claims and the agent tool calls: $7,000 (13%).
A part time statistician on contract: $6,000 (11%).
Compute for the agent sandbox, plus hosting the public dataset: $2,000 (4%).
Subtotal comes to $49,000. We've added a 10% buffer ($4,900) for a total request of $53,900. This amount includes any tax liability for ChatFuse LLC.
The absolute minimum we need is $26,400. That covers the API costs, labels, statistician, and hosting, plus the buffer, but doesn't include my time. If we have to go below that, we'd need to make the human calibration set smaller, and that would hurt the quality of the main finding.
We formed ChatFuse LLC in California in September 2025. ChatFuse is an organization of specialized AI agents, each responsible for a different function (engineering, solutions architecture, QA, research and client communications), led by founder and CEO Nico Coetzee, who would lead this project. The platform is Nico's design and build. That includes the routing layer that picks from over 100 models from places like Anthropic, OpenAI, and Google, the fallback system that retries and switches providers, a Postgres memory setup with vector search, and the tool pipeline that handles moderation and rate limits. We log every model call. We've shipped products on it like cited knowledge assistants, document AI, agent workflows, and speech tools. About 50 people have signed up for the subscription product, which is in early access. On the consulting side, we delivered an AI reporting and operations system that a healthcare marketing agency uses.
We haven't published any research so far, and this dataset would be our first public one. One thing we got wrong: our NSF SBIR Project Pitch from 24 September 2026 didn't get invited because it didn't define one clear high risk technical innovation. We fixed that. Now we structure research around single testable hypotheses, like we did here.
Money: 2025 had $0 revenue (since we started in September). So far in 2026, we've brought in about $87,000, all from consulting. That covers everything; we haven't taken any outside money.
If our attempt doesn't work, grading inconsistencies will probably be the cause. It's tricky to line up individual claims from different answers, and if our hand checked set proves the automatic grader isn't consistent enough, our pairwise scores will have too much spread. That could mean our statistical test won't show a clear result. We'll share those findings regardless, and we'll release the full dataset too. The other big risk is that the model providers push updates, so we're locking every model version and we'll timestamp every single run.
Without this funding we'd do a smaller version later, paid for out of consulting revenue. It would use fewer models and skip the human labeled calibration data. That'd mean results take longer and they wouldn't be as solid.
We're a for profit LLC. All project outputs are released openly; our commercial platform isn't among the deliverables. We haven't published peer reviewed work yet, so a statistician would be brought in for the analysis.
$0. ChatFuse LLC was formed in California in September 2025, has taken no outside money, and has never received a grant.
Also pending, for this exact work: a request to the EA Funds Transformative AI Fund, dated October 2, 2026, for $53,900. If that's funded, our goal here drops by the same sum. No part gets paid twice. Open Parity Bench (a different project of ours, measuring how much energy routed AI systems use and how good their answers are) has its own Manifund page.