Project summary
This project introduces Algebraic Geometric Empirical Process (AGEP) Theory, a groundbreaking unified mathematical paradigm that bridges algebraic geometry (Hartshorne's scheme theory) and infinite-dimensional probability (van der Vaart's empirical processes). Our preliminary work, published with a permanent DOI on Zenodo, establishes the AGEP Theorem: it mathematically proves that by applying a monoidal blow-up (resolution of singularities) to the rank-collapsed determinantal subscheme, the uniform Donsker property (uniform Central Limit Theorem) of the empirical process is algebraically restored in local coordinate charts.
What are this project's goals? How will you achieve them?
The goal of this project is to extend my pilot work on Algebraic Geometric Empirical Process (AGEP) theory from a toy $2 \times 2$ linear attention model to general, high-dimensional, and multi-layer Transformers. Specifically, I aim to solve two major theoretical bottlenecks in deep learning: (1) mathematically proving how the uniform Donsker property (uniform CLT) is restored near higher-dimensional attention singularities using monoidal blow-ups, and (2) formulating Stochastic Gradient Descent (SGD) trajectories as stable, non-degenerate diffusion processes on exceptional divisors (pulling back SDEs onto blown-up spaces).
To achieve this, I will break the research into three technical steps:
Algebraic Formulation: Apply determinantal ring theory (Bruns & Vetter) to analyze the singularity structures of higher-dimensional $d \times d$ attention schemes.
Empirical Process Theory: Construct monomial envelopes on the blown-up affine charts to bound the metric entropy by the Real Log Canonical Threshold (RLCT, $\lambda$) under Aad van der Vaart’s (1996) framework.
Dynamical SDEs: Model the SGD noise covariance matrix on the blown-up space, proving that the Jacobian of the resolution map cancels the degeneracy of the Fisher Information Matrix (FIM).
How will this funding be used?
I am requesting a lean, milestone-based budget of $60,000 USD per year (total $120,000 USD for 24 months) to support this independent research.
The funding will be strictly allocated to:
PI Stipend ($45,000/year): This stipend will allow me to dedicate 100% of my time and intellectual energy to this research, bypassing other consulting work.
Computational Resources ($8,000/year): Cloud GPU instances (e.g., H100 instances on Lambda Labs/RunPod) to run large-scale PyTorch simulations tracking FIM eigenvalues and SGD trajectories in deeper networks.
Workstation & AI Tooling ($3,000/year): Advanced developer environments and API access to sustain my highly optimized human-AI research pipeline.
Outreach & Publication ($4,000/year): Open-access publishing fees and travel expenses to present AGEP theory at top-tier conferences (NeurIPS, COLT, or DevInterp workshops) to gather community feedback.
Who is on your team? What's your track record on similar projects?
I am a solo, independent researcher with a background in mathematics (specializing in topology and mathematical physics). I drive the core conceptual directions, formulate the mathematical hypotheses, and design the theoretical framework.
To overcome the lack of a traditional institutional lab, I utilize a highly optimized human-AI co-working loop: I use advanced LLMs (specifically Gemini Notebook) as an interactive cognitive partner to verify algebraic identities, generate LaTeX code, and write PyTorch simulation scripts under my direct oversight.
My Track Record: In less than 7 days, this lean paradigm successfully produced the foundational AGEP framework. I hand-computed the monoidal blow-up of a $2 \times 2$ attention singularity, verified the resulting "spectral collapse" of the FIM via numerical simulations, and compiled a rigorous 14-page pre-print.
I have published this pre-print with a permanent, citable DOI on Zenodo to secure international priority for this theory:
What are the most likely causes and outcomes if this project fails?
The most likely cause of failure is mathematical complexity. While the monoidal blow-up and normal crossing standard form are highly tractable in my $2 \times 2$ attention pilot study, higher-dimensional determinantal ideals ($d \times d$ multi-head attention) are notoriously complex. The resolution of singularities might yield non-reduced schemes that do not easily admit a single $L_2(P)$ monomial envelope, halting our proof of general Donsker restoration.
Another potential bottleneck is numerical discretization errors. The continuous-time SDE formulation of SGD on the exceptional divisor might prove difficult to simulate accurately in PyTorch due to high-dimensional gradient noise, limiting the empirical validation of our theoretical predictions.
Outcome in case of failure: Even if a general, universal proof remains unsolved at Month 24, this project will still yield highly valuable intermediate results. We will publish the explicit mathematical blow-ups for $3 \times 3$ and $4 \times 4$ attention models, and we will open-source our complete PyTorch codebase for FIM spectral tracking near singularities. This will provide the DevInterp and statistical learning communities with a solid, reproducible dataset to build upon.
How much money have you raised in the last 12 months, and from where?
$0 USD. This research has been entirely self-funded and executed using my own personal computational resources and time. This is my first application for external funding for this project.