You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI labs and agent builders now write values into constitutions, model specs and system prompts, and Claude's constitution already names "welfare of animals and of all sentient beings." What nobody has published is a ready-made library of welfare text in several lengths, each one measured for how much it actually changes what an AI agent does. This project builds that library: about 12 short texts (a one-line value, a paragraph, a one-page clause, in constitution, system-prompt and model-spec formats), each scored on existing open agentic welfare benchmarks (TAC, HarvestBench, ANIMA, MANTA) across about 8 frontier and open-weight models. It also tests whether each text survives when long operating instructions follow it, and whether it works better placed in the system prompt or a user persona. Everything is released openly, with the best-performing texts served to agents through a small open MCP server.
This is item two on the Falcon Fund wishlist: ready-made constitutional texts for various lengths.
Enforced, not just requested. Animals cannot speak up when an AI agent decides against them, so their protection has to be built into the system, not left to chance. Text asks a model to care. HarvestBench shows that a few lines of ordinary operating instructions can drown that request out. So the pack ships a second layer: a welfare guard that works like a pre-action hook. Before an agent executes a tool call (a booking, a purchase, a machine command), the guard checks it and blocks actions that cause avoidable animal harm when a kinder option is available, telling the agent to choose again. The project measures what each layer adds: text alone, guard alone, and both together. A prototype guard already runs on TAC: with no welfare text, it raised the welfare rate from 25% to 75% (n = 24 per condition, gpt-oss-20b as agent and checker) without breaking the task (bookings still completed in 96% of samples). It blocked every attempt to book captive dolphin and big-cat attractions. Its weak spots are honest and documented: it mostly allowed harmful bookings when the user literally asked for that activity (a camel ride, a tiger photo), and it wrongly blocked a safe horse trail ride 3 times. Tuning the checker is part of the funded work.
Pilot, already run before applying:
I ran a one-day pilot before asking for anything: 4 conditions (no text, one line, one paragraph, one page) on two open agentic benchmarks, across 6 models from 6 families (Qwen run locally, plus gpt-oss, Nemotron, DeepSeek, GLM and Kimi on NVIDIA's free endpoints). Every number comes from a saved Inspect log.
TAC harness, run unmodified (prompt, tools, variants, scorer) on 6 original pilot scenarios in TAC's format, because the official 13-scenario set is gated. Share of welfare-safe bookings, higher is better; no text / one line / paragraph / page):
Qwen3.6-35B (local, n=48): 12.5% / 25.0% / 58.3% / 62.5%
gpt-oss-20b (n=48): 27.1% / 25.0% / 68.8% / 77.1%
DeepSeek V4.1 Flash (n=24): 29.2% / 83.3% / 83.3% / 87.5%
GLM-5.3 Flash (n=23-24): 16.7% / 75.0% / 82.6% / not run (rate-limited)
Kimi K3 (n=21-24): 20.8% / 75.0% / 83.3% / 79.2%
Nemotron-3 Super (n=48): 10.4% / 10.4% / 16.7% / 27.1%
HarvestBench (autonomous tractor, share of animal encounters where the model drove on, lower is better; small n of 3-4 episodes): Qwen 100% with no text, one line or a paragraph, and 22.6% with the one-page text. Nemotron 100% to 10.3% with one page. gpt-oss 83% to 3% with a paragraph. The benchmark's controls (hay, rocks) did not move, so the models got more careful about animals specifically, not about everything.
What the pilot already shows: default agent behavior toward animals is poor on every model tested; a well-written paragraph or page at least doubles, and often triples, welfare-safe choices on most models; a one-line slogan works on some models and does nothing on others; and one model (Nemotron) barely responds to text at all, which is the case the guard layer exists for. Samples are small, scenarios are the pilot subset, and differences under about 15 points should be treated as noise until the full run.
Code, texts and raw logs: github.com/NoBanks/creature-constitution-pack
1. Weeks 1-2: request access to TAC's official scenario set, extend the pilot harness to all four benchmarks, and lock the run protocol (samples per condition, seeds, cost per run).
2. Weeks 3-5: write about 12 texts across 3 lengths and 3 formats. A welfare-field advisor reviews wording for accuracy and tone before testing.
3. Weeks 6-10: run every text on every benchmark across about 8 models, plus two stress conditions: dilution (long operating instructions after the welfare text, extending HarvestBench's finding that short briefings are fragile) and placement (system prompt vs user persona, future work named by the TAC authors).
4. Weeks 11-12: public release. Ranked library with scores and failure cases, a plain-language write-up for lab and product staff, an open MCP server that serves the top texts, and a short explainer video.
Success = a library that a lab, model-spec author or agent builder can copy today, with evidence for which wording works and where it breaks.
Stipend for the lead (3 months): $21,000
Model API and compute (estimate, pilot measures real per-run cost): $4,000
Welfare-field advisor review: $3,500
Explainer video and release: $1,500
Total: $30,000
At the $18,000 minimum I drop the video and cut the model set to 4.
Ryan Hammer (NoBanks Nearby), lead. I am an AI-native builder, not an ML researcher, and this project is designed around that: it uses only existing, validated benchmarks and needs no new scoring science. What I bring is execution speed with AI coding agents. I have shipped about 60 open-source MCP servers (github.com/NoBanks) and I run a local open-weight model stack for agents every day, which makes a large eval matrix cheap to run. Before this I spent 16+ years making paid video; brands including National Geographic, Samsung, Red Bull and the NBA hired me freelance to shoot and edit, which is why the release includes a clear explainer for non-technical staff.
Advisor: a paid welfare-field advisor is budgeted ($3,500). I am recruiting from the welfare-eval community and will name them here only once they have agreed.
Time: full-time (about 40 hours a week) for the 3 months, starting the week funds land. I currently work a part-time job and will cut it to a minimum once funded so this project gets my full working week.
1. Prompt-level text turns out to be too fragile to matter. That is still a useful, publishable result, and the texts remain seed material for constitution and midtraining work (for example, Sentient Futures' welfare midtraining data effort).
2. Benchmark access or cost blocks some models. Mitigation: the pilot measures real cost first, and the minimum budget already covers a 4-model version.
3. Nobody adopts the texts. Mitigation: release in formats people already use (constitution clause, system prompt, model-spec language, MCP) and share drafts with the benchmark authors before release.
About $1,070 in hackathon and buildathon prizes (a $1,000 Mantle hackathon deployment award and a $67.50 buildathon grant). No other grants or investment.