@Richard I agree that hands-on experimentation is essential. One challenge I've repeatedly encountered while evaluating frontier AI models is that many interesting failures are observed once but aren't preserved in a way that enables independent verification or follow-up research. I'm currently developing an evidence-based methodology focused on documenting, preserving, and reproducing AI evaluation evidence. Since you mentioned information aggregation and verification as an interest, I'd be interested to know whether you think standardized evidence preservation could become a useful part of AI safety evaluation infrastructure.