Koshiro Aoki

master’s student in Computer Science, Waseda University

View PDF

Research Interests

Metacognition | Introspection | (Mechanistic) interpretability | NeuroAI

Education

Waseda University

M.S. in Computer Science
Adviser: Daisuke Kawahara

Apr. 2026 - Present
Tokyo, Japan

Waseda University

B.S. in Computer Science
GPA: 3.8/4. Dean's Award.

Apr. 2022 - Mar. 2026
Tokyo, Japan

Work Experience

Waseda University

Research Assistant
Research on differences between the internal representations of the human brain and language models as part of a JST CREST project.

Apr. 2026 - Present
Tokyo, Japan

The University of Tokyo

Research Contractor
Research on interpretability and defenses of adversarial attacks on robot foundation models.

May 2025 - Present
Tokyo, Japan

Research and Development Center for Large Language Models, National Institute of Informatics (NII)

Technical Assistant
Research on bottom-up interpretability of training dynamics in large language models.

May 2025 - Mar. 2026
Tokyo, Japan

Publications

Conference

Bottom-Up Interpretability of Pretraining Dynamics via Loss Curve Decomposition

Koshiro Aoki, Masaru Isonuma, Max Müller-Eberstein, Yusuke Oda, Hirokazu Kiyomaru, Takashi Kodama, Chaoran Liu, Yohei Oseki, Yusuke Miyao, Daisuke Kawahara

Third Conference on Language Modeling (COLM 2026)

2026

Workshop

Crosscoders Identify Shared or Specific Features between the Human Brain and Language Models

Koshiro Aoki, Itsuki Hamada, Naho Orita, Daisuke Kawahara, Hiromu Sakai

Mechanistic Interpretability Workshop @ ICML 2026

2026

In-Context Neurofeedback: Can Large Language Models Control Their Internal Representations through Privileged Access?

Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Yusuke Haruki, Daisuke Kawahara

Trustworthy AI for Good (AI4GOOD) Workshop @ ICML 2026 Spotlight

2026

Evaluating the Impact of SAE-based Language Steering on LLM Performance

Sebastian Zwirner, Wentao Hu, Koshiro Aoki, Daisuke Kawahara

EACL 2026 Student Research Workshop

2026

Testing Simulation Theory in LLMs' Theory of Mind

Koshiro Aoki, Daisuke Kawahara

IJCNLP-AACL 2025 Student Research Workshop Oral

2025

Domestic Journal

Mechanistic Interpretability: A New Trend in Interpretability Research

Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Hakaze Cho

Transactions of the Japanese Society for Artificial Intelligence Best Paper Award

2026

Honors and Awards

Spotlight Paper

Trustworthy AI for Good (AI4GOOD) Workshop @ ICML 2026

2026

Best Paper Award

The Japanese Society for Artificial Intelligence 40th Anniversary Commemorative Paper Award

2026

Dean's Award for Excellence

School of Fundamental Science and Engineering, Waseda University

2026

Young Researcher Encouragement Award

The 32nd Annual Meeting of the Association for Natural Language Processing

2026

Oral Presentation

IJCNLP-AACL 2025 Student Research Workshop

2025

Encouragement Award

The 20th Young Researcher Association for NLP Studies

2025

Sponsored Award

The 31st Annual Meeting of the Association for Natural Language Processing

2025