AI safety research based in Tokyo.
We are a research team focused on AI safety, alignment, and governance. Our work spans evolutionary approaches, multi-agent coordination, and building interpretable AI systems.
Meet the people
35 researchers and engineers behind the work.


































Research Areas
Pushing the boundaries of technical AI safety through alignment, interpretability, and governance research.
Alignment & Coordination
Ensuring AI systems (individually and in groups) reliably understand and act on human intentions.
Bias & Persona
Understanding the biases language models absorb and the personas they adopt, and how both shape model behavior.
Governance
Studying the policy, legal, and institutional frameworks needed to govern AI safely and responsibly.
Interpretability
Reverse-engineering the internal mechanisms of neural networks to predict, verify, and steer their behavior.
Metacognition & Hallucination
Investigating whether models know what they know, and why they confidently state things that are false.
Security & Robustness
Probing and hardening AI systems against adversaries — from prompt injection and red-teaming to adversarial robustness.




