This is a grant covering work Celeste did as part of MATS with me (during the period pre MATS where MATS isn't able to cover stipends). I think Celeste did great work here, most notably working on a forthcoming paper I really like training activation oracle variants that verbalize the contents of a LoRA to explain the difference between two models. I have an obvious conflict of interest here, but I think this is clearly above the bar.