You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Our aim is to use AI to extract claims from scientific data and help experts reason over them. Our first target is to make this useful for the scientific community. If our results are solid, we will expand this tool for journalists and policy analysts as well. This project first launched in June 2026, though we had been thinking about it for some time longer. The opportunity came when one of us was pursuing a BlueDot course on AI Safety and learned about an AI epistemics competition run by an AI Safety foundation, which gave us the right framing to take the problem seriously. Today we have an early prototype we want to refine.
We are taking small and measurable steps towards this ambitious goal. For instance, Google recently published work on verifiable research pipelines built on evidence chains, which tackles the same problem (Link: https://research.google/blog/science-one-framework-a-verifiable-autonomous-research-framework-via-chain-of-evidence/).
Use AI to automatically extract relevant scientific claims from large collections of research articles.
Evaluate every module and every result against pre-registered criteria and metrics.
Maintain auditability and traceability as a design principle. Any result the user sees is traceable to its exact source passage.
Pay our software engineer to dedicate more time to the project. They currently work full-time at another company to finance this work, and are putting 40+ hours per week into this project on top of that job and their personal life (if any is left).
Computational resources for running experiments at scale: API credits (OpenAI, Anthropic, and similar providers), and storage servers.
Pay human experts to audit our tool and provide feedback to iterate and improve the solution.
Pay our team for writing research articles and publishing results on the relevant topics, such as AI, AI evaluation, methodology, semantics and ontologies, philosophy of science and philosophy of AI.
Conference attendance, if our contributions are accepted.
Exploratory budget:
Fine-tuning an LLM, if our evaluation shows it is needed.
GPU- powered infrastructure, if needed later.
Project lead: independent researcher in AI and Natural Language Processing, with a bachelor's degree and three master's degrees. I have contributed to developing an AI automation framework for cybersecurity and published the results, and I have worked on NLP research, including annotation guidelines, NLP-based solutions, and AI research. I designed and built the current prototype.
Philosophy of Science PhD: supports methodology and epistemology, discusses and stress-tests ideas, co-writes research articles, and serves as an expert evaluator and annotator in their speciality, which is one of our use cases.
Biology PhD (male fertility): supports methodology, discusses ideas, co-writes research articles, and serves as an expert evaluator and annotator in their speciality, which is another use case.
Most likely causes:
Lack of financial resources to buy more time. Most of the team cannot dedicate themselves full-time yet.
Difficulty of the task. Auditable information retrieval over scientific text is among the hardest problems in natural language processing. The same finding can be phrased in many ways, and the same terms can mean different things across fields. This is hard to do without LLMs, and relying on them at scale makes the task expensive.
Need for more team members willing to push this problem forward together. We already have a mathematician and a physicist interested, but we will need to advance the prototype first.
Outcomes in case of failure:
Circularity in the audit instrument. We will evaluate different LLMs across tasks. We will seek to avoid confounding and isolate variables. Models from the same family may share the same failure modes, and LLMs may not perform as expected on this task, which places a ceiling on how far this problem can be resolved.
Precision-recall trade-offs in comparability instruments. We are exploring trustworthy configurations for retrieving comparable information, and we have already encountered precision-recall trade-offs that depend on which instrument is used.
Insufficient gold-standard annotation, or a highly complex task to annotate reliably. It is possible that the task does not reach a minimum inter-annotator agreement among human experts.
Cross-domain generalization, as our current focus is biomedicine and health sciences. Our second use case is philosophy of science, which is a deliberately extreme one where subjectivity is much harder to handle, and results may not transfer.
Limited team time for doing research.
Everything we produce along the way remains usable: published papers, accepted conference contributions, processed datasets and claim graphs, expert annotations. Negative results are informative in their own right, and we think that documenting this process and understanding why an instrument fails is itself a contribution.
$500 (minimum): covers basic computational resources to keep running experiments. Beyond money itself, reaching the minimum tells us that someone out there cares about this problem, and that motivation matters to us as well.
Up to $20,000 (full) and $50,000: buys dedicated research and engineering time, expert annotation, conference attendance and publication costs, and servers for scaling and storing processed data.
More than $50,000: buys dedicated research time for more team members and gives us the financial stability to commit fully to this project.
Anything in between moves us proportionally and we are very grateful for every contribution.
Thank you for your attention and support!
Have a nice day and kind regards.
We have won small amounts of computing resources in two competitions: $145 in one and $200 in AWS credits from another. We are not naming the competitions here because both programs are still ongoing and one involves anonymous review of our submission. We will update this section once results are public.