You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project Description
An AI-generated description can sound convincing and still give a student the wrong information. It might describe the wrong trend in a graph, omit an important qualification, or change a value while repairing a table. Making content easier to access must also preserve what it says.
I’m seeking up to US$20,000 for an open evaluation project: test when AI-assisted accessibility remediation preserves the source’s meaning, and when the workflow should stop for human review. The outputs will be a public test set, evaluation software, independently reviewed findings and a guide for reproducing the work.
Over the last ten months, I have built and self-funded Aelira. Its open-source core is public and in beta, with document scanning, supported remediation and review workflows. This gives the study an existing implementation to work with, while its reliability still needs evaluation.
Code and setup: https://github.com/Aelira-AI/aelira-core
Hosted demo: https://aelira.ai/us/demo
The results would help accessibility teams compare workflows and review effort, and give developers a reusable way to test whether changes introduce errors. Everything funded here will be available independently of Aelira’s paid hosting.
What are this project’s goals? How will you achieve them?
The main question is whether checking outputs against source evidence reduces errors enough to justify the extra cost and review effort.
At full funding, I will target 150 rights-cleared cases covering image descriptions, chart interpretation and supported table repairs. Cases will include missing evidence and embedded instructions that could distract a model. Thirty STEM cases will be held back while prompts and checks are developed. Confidential university documents and student records are outside the study.
Image-description tasks will provide the actual image. Deliberately missing-evidence cases will test whether the system pauses rather than inventing content.
I will compare three approaches on the same cases:
- Ordinary prompting.
- Prompts explicitly requiring preservation of the source’s meaning.
- Source-evidence checks with escalation to human review.
Four model configurations, three approaches and three runs per case give up to 5,400 outputs. I will check compatibility and costs first, then hold configurations fixed for the comparison.
Across all runs, the software will record completion, refusals, latency and cost, and compare reference facts where exact checks are possible. Independent reviewers will assess factual errors, omissions, accessibility usefulness and review time on a sample of 200 outputs. A second reviewer will score 50 of those to measure agreement. Reviewers will see the source material with model and method labels hidden; factual fidelity and accessibility usefulness will be scored separately.
The sampling procedure and rubric will be published before the main runs, covering task types, models and approaches. The analysis will account for repeated runs of the same case. Human judgments will apply to the reviewed sample, with uncertainty and limits reported explicitly.
The twelve-week plan is:
- Weeks 1–2: agree case rights, scoring, sampling and reviewer arrangements; run a timed workload pilot.
- Weeks 3–5: build the case set and evaluation software.
- Weeks 6–9: run comparisons and independent review.
- Weeks 10–12: analyse results, release artifacts and document reproduction.
I will release the standalone harness under MIT, newly authored cases and labels under CC BY 4.0 where rights permit, and core improvements under its existing AGPL-3.0 licence. Releases will include configurations, model versions, costs and negative results. Any changes to the core will follow the evidence from the study.
This is an exploratory comparison using existing models. It does not promise a new trained model or certify accessibility conformance.
How will this funding be used?
Funding buys dedicated research time and independent scrutiny. I have already purchased a Founders Edition DGX Spark for testing; it is available to the project. The budgets below fund future work.
All amounts are US dollars. Rates are planning estimates, with reviewer availability and quotes still to be confirmed.
US$5,000 — six-week feasibility study
- $2,000: research and engineering, 50 hours at $40/hour.
- $2,000: independent accessibility and technical review, 20 hours at $100/hour.
- $500: model APIs and supplementary compute.
- $250: hosting, storage and research artifacts.
- $250: independent release and reproducibility review, 2.5 hours at $100/hour.
This delivers 20 image/chart cases, two model configurations and two approaches, with three runs each: up to 240 outputs, 40 first reviews and 10 second reviews. I will release the small test set, working harness and findings. It funds about 8.3 hours of my time per week.
US$10,000 — eight-week study
- $4,500: research and engineering, 112.5 hours at $40/hour.
- $3,500: independent accessibility and technical review, 35 hours at $100/hour.
- $1,000: model APIs and supplementary compute.
- $500: hosting, storage and research artifacts.
- $500: independent release and reproducibility review, five hours at $100/hour.
This expands to 60 cases, three models and three approaches, with three runs each: up to 1,620 outputs, 100 first reviews and 25 second reviews. It funds about 14.1 hours of my time per week.
US$20,000 — full twelve-week study
- $9,000: research and engineering, 225 hours at $40/hour.
- $7,000: independent accessibility and technical review, 70 hours at $100/hour.
- $2,000: model APIs and supplementary compute.
- $1,000: hosting, storage and research artifacts.
- $1,000: independent release and reproducibility review, ten hours at $100/hour.
This funds the full study above and about 18.75 hours of my time per week. My work covers cases, software, experiments, analysis, documentation and relevant core changes. Independent release review is separate external work.
I have bootstrapped this as far as I can. I’m a full-time carer and run a small IT consultancy alongside Aelira. External support would let me protect research time and pay people qualified to challenge my work, rather than continuing to carry every cost myself.
Who is on your team? What’s your track record on similar projects?
I’m Reginald Crampton, Aelira’s founder, project lead and company Director. I am its sole technical maintainer, handling development, infrastructure, security and model testing. I have also contracted with Outlier and CrowdGen on model evaluation, red-teaming and safety work. That work is under NDA, so I cannot publish client details or results.
My co-founder Erik Vuchich leads external partnerships, sales and outreach. We are both based in Australia. I will lead the study; Erik’s outreach role can support reviewer recruitment and sharing the public outputs. Funded lead research hours are allocated to me.
The public core and its documentation are the evidence of what I have built. They are a starting point for this research, not proof that generated content is reliable. I will recruit reviewers unaffiliated with Aelira who have relevant accessibility and technical expertise. Recruitment is pending, and they will be paid regardless of their findings. My commercial interest will be disclosed.
Longer term, I’m exploring a separate accessibility-AI research foundation using aelira.org, which I own. It is a future option, not an established charity or the recipient proposed for this grant.
What are the most likely causes and outcomes if this project fails?
The main delivery risks are my limited time, recruiting suitable reviewers, and a workload or compute bill larger than expected. The timed pilot will check these assumptions before the main study. If the planned scope does not fit the budget, I will agree changes with the funder before proceeding.
The safeguards may not help. They could reject useful tasks, add expense or move too much work onto reviewers. Results may vary between tasks and fail to generalise. I will publish those findings alongside positive results; identifying where an approach fails is part of the purpose.
If I cannot complete agreed work, I will notify Manifund and donors promptly and agree how to handle reduced scope and unspent funds.
Without funding, I will need to spend more time on paid consultancy. The open core will remain available, but the independently reviewed research will slow down or stall.
How much money have you raised in the last 12 months, and from where?
US$0 in external funding. Hardware, hosting, model access and running costs have come out of my own pocket.
I have also submitted applications or enquiries to BlueDot, Sentient, GitHub and Paradigm 3. Emergent Ventures declined my application. No funding has been awarded. I will disclose any awards and remove overlapping costs or agree distinct additional work before spending funds.