You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This project attempts to observe how an LLM (Large Language Model) behaves when self-referential/ego vectors are fed into it using activation steering, and its attention matrix is observed to check whether it simulates a non-dual-like state.
I am interested in observing if I can induce a non-dual-like state (similar to ego dissolution) in a large language model using activation steering. I expect to see measurable changes in the model's attention entropy. Specifically, the attention should become much more globally distributed rather than concentrated at a certain point. This is basically a way to computationally test a specific prediction from the Dual-Aspect Monist (DAM) framework of consciousness which I've been working on(the manuscript of which can be found here: https://philpapers.org/rec/KHUTMO ).
I'll be working with either GPT-2 or Llama. I plan to take a self-referential vector and inject it into the model at different intensities using activation steering. After that, I will measure the attention entropy across the model's layers. I'm trying to observe if the entropy hits a maximum at an intermediate ego density. If it does, that could possibly validate the DAM prediction that dissolving the ego leads to a highly integrated, global state of awareness.
I've already built and run four experiments leading up to this. I've done self-referential clustering (using cosine similarity and PCA), a layer-by-layer analysis where I pinpointed Layer 8 as the point of ego-separation, some circuit analysis on attention heads, and an IIT Phi sweep that actually showed an inverted-U relationship with ego density. The previous experiments can be found here: https://github.com/oxerz8/Ego-Density-Activation-Steering-and-Integrated-Information
These experiments have been reviewed informally by leading researchers in mathematical consciousness, including a co-author of the IIT 4.0 framework. Based on this work, I have been invited to collaborate informally with an IIT researcher over the next year to transition into a formal MSc program in Fall 2027. This grant will fund my independent work during this gap year.
The main risk is that the attention entropy metric is not sensitive enough to catch the effect I'm trying to observe. It is also possible that the effect shows up in one model but doesn't scale or replicate across different model sizes. But even a negative result is useful here, because it helps map out the limits of what LLM internals can actually tell us about theories of consciousness.
0
There are no bids on this project.