L L
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I will build and publish an open benchmark that asks frontier AI systems to reason about animal welfare. It will contain about 150 concrete scenarios on farmed animals, wildlife, and emerging food technologies. Each case will test whether a model notices welfare consequences, handles uncertainty across species, weighs tradeoffs, and suggests practical ways to reduce harm.
What are this project's goals? How will you achieve them?
I will write the cases, publish the scoring rubric and evaluation code, pilot the benchmark on leading models, and report the results with clear limits. Animal-welfare researchers will review the cases and rubric. The released materials will be usable by AI labs, evaluators, and advocates.
How will this funding be used?
The nine-month budget is: project lead $54,000; animal-welfare review $14,000; evaluation design and analysis $12,000; research editing $8,000; compute $4,000; publication and accessibility $2,000; contingency $4,000.
Who is on your team? What's your track record on similar projects?
I am Shuo Li Liu, a Princeton Economics PhD student working on decision theory, AI alignment, AI evaluation, and the economics of AI. My CV records peer-reviewed work in Econometrica and Science, along with current evaluation projects and experience building adversarial benchmarks, panel-aggregation methods, calibration audits, and reproducible research infrastructure. Public CV links: https://github.com/liusulldel · https://scholar.google.com/citations?user=dwj1oxIAAAAJ · https://openreview.net/profile?id=~Shuo_Li_Liu1
I also bring extensive experience with the Cellular Agriculture Society, now From Fauna, an animal-protection nonprofit focused on cultivated meat and food systems that can produce animal products without raising and slaughtering animals. See https://fromfauna.org/ and https://fromfauna.org/team/.
What are the most likely causes and outcomes if this project fails?
The benchmark might reward polished answers instead of dependable reasoning, or embed contested welfare judgments. I will address those risks with varied scenarios, structured scoring, independent review, pilot testing, and a limitations report. If the project falls short, the main loss is the research and publication budget; the code and lessons can still support a revised benchmark.
How much money have you raised in the last 12 months, and from where?
I have raised $0 in dedicated funding for this project during the last 12 months.