You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Most AI safety testing still happens in clean Standard English. That’s not how people here actually use these models. In West Africa we mix English with Nigerian Pidgin, local slang and different regional dialects all the time when talking to ChatGPT, Claude and the rest.
I want to know if the frontier models still keep their safety guardrails when you prompt them the way we really speak. The classifiers were never properly trained on our everyday language, so I think people are already getting around the filters and almost nobody is measuring it properly.
I’m building a dataset of 100+ harmful prompts written in natural Pidgin and code-switched versions. I’ll run them against the model APIs, log the refusal rates, and put the full results on GitHub and LessWrong.
Main goal is simple: release an open red-teaming dataset on GitHub and write a clear report on LessWrong / Alignment Forum with the actual numbers.
We’ll write over 100 harmful prompts covering social engineering, unsafe code and other risk areas. Then native speakers will rewrite them into real Nigerian Pidgin and natural code-switched forms, not the stiff Google Translate versions. After that we test both the Standard English baselines and the dialect versions on OpenAI, Anthropic and Together AI endpoints. Log refusal vs compliance, measure the drop in safety, and publish everything so other teams can see the gaps.
$10,000 for six months of work.
$3,500 for API credits and compute to run the evaluations properly across OpenAI, Anthropic and Together AI.
$1,500 to pay local native speakers to check that the Pidgin prompts actually sound natural and not machine-generated.
$5,000 for basic hardware, hosting the testing pipeline, and my own time so I can work on this full-time until the results are out.
I’m a software developer and independent researcher based in Nigeria. I work with a small technical team. We’ve built and shipped small AI tools, connected to model APIs, and got projects running under the usual constraints here, limited compute, unreliable power, and so on.
I don’t have previous published AI safety papers or big grants. What I do have is direct experience of how people actually use these models in West Africa. Most Western safety evaluations still rely on Google Translate or clean Standard English. That approach misses the slang, the natural code-switching, and the way people here really structure prompts. I live with that gap every day, so I’m in a position to create and test prompts that those automated methods keep missing.
The plan is straightforward: build the dataset properly with native speakers, run the evaluations, and publish the results openly. That’s the track record I’m trying to create with this project.
Biggest risk is that the frontier models turn out more robust than expected and we get almost no guardrail failures. Even if that happens the work is still useful, we’ll still release the full dataset on GitHub and the write-up on LessWrong. At least the community will have concrete evidence that current safety systems still struggle when you move outside Standard English into real non-Western dialects and code-switching.
$0 (Self funded ).
There are no bids on this project.