- 🗾 JCA-Bench — Japanese Cultural Appropriateness Benchmark for evaluating AI cultural understanding
- 🔬 Competing Circuits Across Languages — Safety vs. Instruction-Following Dynamics in Multilingual LLMs (WIP)
- 📚 LLM Security, Alignment & Governance Resources Awesome — Curated list of LLM security, alignment & governance resources (maintained as 0xSweet)
- 🎭 Gandalf Agent Breaker Report - 2nd place in model league
- Mechanistic Interpretability
- Multilingual AI Safety
- Red Teaming