What this project is
Biological AI systems are increasingly being connected to external scientific databases. These systems do not only answer from the information that is contained in a model but also retrieve protein sequences, cell profiles, gene annotations and experimental results before recommending an action to the user. This retrieval can make biological AI a lot more accurate but also creates a huge risk. The model can treat retrieved information as 100% trusted evidence even when that information is incorrect, compromised or even deliberately manipulated!
In this project we will develop BioRAG-Guard. An open-source benchmark and defense framework which evaluates data-poisoning risks in biological AI systems. Can a small number of malicious or corrupted database records influence what a biological AI agent retrieves and change the agent's conclusions? This is the main research question of the project.
We will evaluate different representative biological retrieval architectures. More specifically, protein embedding retrieval based on models such as ESM-2, single-cell similarity search using SCimilarity, reference mapping and nearest-neighbor label transfer using scArches and finally retrieval-augmented prediction of cellular perturbation responses. These are systems span across different biological data and different ways in which scientific agents rely on external evidence.
For each of the systems we will:
1) Construct controlled and non-hazardous poisoning scenarios using public or synthetic data. The evaluation will measure whether manipulated records a)entered the highest ranked retrieval results b)change a predicted label or a biological interpretation c) transfer across embedding models / evade quality control procedures. The benchmark will distinguish between retrieval failure (corrupted information returned) or an outcome failure (corrupted information changes the answer of the agent).
2)Implement practical defenses. These will include source and dataset provenance, visioned by database snapshots, anomaly detection across metadata /embeddings, quarantine periods for recently added records. Conclusions must be validated by multiple independent records.
The main outputs of the project will be an open-source evaluation that will include a collection of safe and reproducible poisoning scenarios, baseline results across multiple biological retrieval systems and implementation of several different defenses. Additionally, we will supply documentation for developers and database maintainers and an academic paper of the projected.
The project will NOT involve a) pathogen engineering, harmful biological sequence design or wet lab experimentation.
https://app.grantmaking.ai/projects/9f03beeb-bcca-441b-bfae-b1339a96ed0b
Theory of impact
Future biological AI systems may move beyond answering scientific questions and begin autonomously retrieving evidence, comparing biological sequences, classifying experimental samples, prioritizing hypotheses, and recommending laboratory actions. As these systems become more capable, their behavior will depend not only on the underlying AI model but also on the integrity of the databases, vector indexes, and reference collections from which they retrieve information.
A well-trained and well-aligned model can still produce dangerous or misleading outcomes if the evidence presented to it has been manipulated. A corrupted protein annotation, single-cell reference profile, or perturbation record could influence an agent’s reasoning while appearing to be legitimate scientific evidence. This creates a potential pathway through which an external attacker, compromised data source, or systematic database error could redirect the behavior of a more autonomous biological AI system.
Most current biological AI safety work focuses on the capabilities and outputs of the model itself. Examples include whether a model provides harmful biological knowledge, whether it follows unsafe requests, and whether hazardous capabilities can be removed or controlled. These are important questions, but they do not fully address the retrieval layer. As scientific agents become more dependent on external tools and databases, retrieval integrity becomes part of the safety boundary.
BioRAG-Guard will reduce this risk by creating a repeatable way to test biological AI systems under controlled data corruption before they are deployed in higher-consequence settings. The benchmark will show which architectures are vulnerable, how many manipulated records are required to change a result, which failures are detectable, and which defenses preserve normal scientific performance while reducing manipulation success.
The immediate impact will be to provide developers with concrete evidence about an underexamined failure mode and practical defenses they can integrate into their systems. Biological database maintainers may also use the findings to improve provenance tracking, ingestion controls, versioning, anomaly detection, and review procedures.
The longer term goal is to establish retrieval and data integrity as a standard component of biological AI safety evaluations. If future biological agents are tested not only for harmful capabilities but also for their behavior when their evidence sources are compromised, dangerous failures may be identified before these systems gain substantial autonomy or are connected to real laboratory infrastructure. This project addresses one specific but potentially important layer of biological AI risk: preventing capable systems from being steered through the scientific information they are taught to trust.
How the money will be spent
Minimum budget:
$21,000: Researcher support and protected research time for system implementation, benchmark design, experiments, analysis, documentation, and preparation of an academic paper.
$5,000: Computing infrastructure such as GPU access, cloud and storage costs, HPC environment.
$2,000: Review by researchers with relevant biological backgrounds.
Ideal budget:
$4,000: Part time research engineering student for experiment automation.
$2,000: Reproducibility and security review, documentation, verification that benchmarks can run independently.
$2,000: Dissemination and maintenance, including publication costs/ conference participation
The team will mainly consist of researchers from Prof. Ilias Georgakopoulos-Soares's lab at UT Austin.
1)Michail Patsakis
2)Kimon Antonios Provatas
3)Charalampos Koilakos
4)Aris Karatzikos
4)Christos Galanopoulos
The main risk is that evaluating several specialised systems takes longer than expected. In that case, I will prioritise fewer systems while preserving the core benchmark and defence evaluation. The attacks that will be tested might also fail to change downstream outcomes and agent behaviour. This would still provide useful evidence about which biological retrieval architectures are naturally robust.
BioRAG-Guard has been recommended for $38,000 from the grantmaking.ai Launch Round. This project has not received any other funding.