You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I got interested in this problem after using AI systems for long periods in ordinary day-to-day work. Over time I noticed that information remembered from earlier conversations can keep affecting later responses. Sometimes that is useful, but if the memory contains an error, outdated information, or a contradiction, that problem can also persist.
I want to test this properly with experiments instead of relying only on my own observations. I want to see whether corrections survive long breaks, what happens when memories conflict, and how the same memory behaves when used with different models.
At the same time, I want to build a small working memory prototype where I can see where a memory came from, how it changed, what was corrected, and roll back the state when something goes wrong.
I do not know in advance how strong the effect will be. If it turns out to be small or highly dependent on the specific model, that is still a useful result.
The funding would mainly let me work on this project properly for nine months instead of doing it only in spare time. The main costs are my research time, model and API access for a large number of experiments, compute, data storage, and backups.
Part of the budget is for a research workstation and technical infrastructure, because I want to store and process the research data reliably instead of depending only on cloud services.
Some funding will also go toward software tools, repeated experiments, and external technical or methodological review of the results.
The full project budget is $44,000 for nine months.
I am currently working on this project alone as an independent researcher. I do not have a separate research team at this stage.
I have already built a working experimental process and completed several test series. The main study included 72 runs with a pre-defined plan, automated response collection, and separate scoring of the results. I then completed another 24 calibration runs to make the methodology more sensitive.
The first study produced a null result: there were no unsafe decisions in any condition. I did not try to force a stronger conclusion. I kept the result, documented the design limitation, and changed the methodology.
That is the main experience I bring into this project: freezing conditions in advance, preserving raw data, keeping inconvenient results, and changing the method when an experiment exposes a weakness.
The most likely way this project could fail is that the effects of long-term memory turn out to be weak, unstable, or highly dependent on the specific model and memory architecture. In that case, it may be difficult to reach a general conclusion that transfers across systems.
Another risk is technical limitations in APIs and model behavior. Some comparisons between systems, or transferring the same structured memory between models, may be harder than expected.
If the main hypothesis is not supported, I would still consider the project useful if it becomes clearer where the problem does not reproduce and which measurement methods are unreliable.
In the worst case, I would still end up with a working prototype, a set of reproducible tests, and well-documented negative results rather than a strong general effect. That would still provide a useful basis for future work and help other researchers avoid repeating the same mistakes.
I have not raised any external funding in the last 12 months. The preliminary work in this area has been self-funded so far.
I also currently have a pending application with EA Funds / the Transformative AI Fund for the same nine-month project with a $44,000 budget. If I receive funding from another source, I will update Manifund and avoid double-funding the same expenses.
Before this, I submitted a separate EA Funds application for a broader 12-month project requesting $200,000. That application was declined and no funding was received.