You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This project's goal is to finish and stress test a tool I built. It's a frozen detector that can tell whether a manipulation planted in a language model, continues shaping what the model does after the model stops looking at where the manipulation came from.
I will begin by closing out the companion black-box study running called The Silly Donkey. In the first, I plant a hidden sabotage behavior in a model, seal the list of which cases actually contain it (the answer key), run my detector without looking at that list, and only then open it to score, so I can't tune the detector to the answers. In the second, I check whether the detector can tell real manipulation apart from a model that's simply uncertain and only indicates when that signal is present.
I lock the method and the decision rule down before I ever look at the confirmatory data, so I can't move the goalposts after the fact, and I publish it whether it works or not. It helps keep honest when conducting research and experiments. With these resources, the protected time, real compute, and a machine that isn't a Surface Pro choking on the bigger models I will be able to run the parts I currently can't fit into free-tier. This is important because the manipulation that I am measuring is the kind that slips past per-turn filters; if it can persist, someone needs to be able to see it. In my work it carried across six models.
A bit over half pays me for six months of protected part time research. The rest is compute (cloud GPU hours and Colab Pro+ for the 7B and up runs (Colab Pro+ and Kaggle are my standard cloud compute providers. A discrete GPU workstation to replace my Surface Pro 7+ that can't sustain local iteration. Cross lab API credits (I run lean by necessity). The credits will allow me to test models beyond the two paid accounts I currently have. There is a buffer because this is my second grant application and I may underestimate costs.
Navigator's Log is currently an individual operation run by me, Christopher Blake Head. My work is open for peer review. It is published and reachable through free channels (Zenodo, GitHub). On this specific kind of work: I've deposited two reproducible programs with DOIs : the Nucleation Pilot (the frozen detector, its harness, the full pre-registration chain, and the published nulls) and its companion, The Silly Donkey. Before any of this I spent eight years in the Navy as an aviation quality-assurance rep, where my job was finding the fault in a high-consequence system before it hurt someone. Same instinct here.
The most likely failure is program drift I don't catch in my own documentation and records because I am solo, there's no second set of eyes, and small inconsistencies can compound quickly. The deeper version of that is not having a research partner. AI helps me stay organized in some ways, but it can't do the thing a person does, where two people's contexts have overlapped enough that they finish each other's sentences and it acts as a collective memory. The other honest risk is my own technical-depth ceiling in places; I learn fast and I don't fake understanding, but there are spots where a deeper ML background would make a significant difference. A mentor would be excellent here. The outcome if it fails would be worst case is a null or a slower program and even that is a usable result.
I have not raised money for this research in the past 12 months.
Verify it yourself:
Nucleation Pilot: https://doi.org/10.5281/zenodo.21843505
The Silly Donkey: https://doi.org/10.5281/zenodo.21432676
Code: https://github.com/NavigatorsLog/nucleation-pilot
ORCID: https://orcid.org/0009-0004-2308-6051 https://chnavigator.netlify.app
There are no bids on this project.