Project

Cross-Lingual Concept Steering and Interpretability in Language Models

Cross-Lingual Concept Steering and Interpretability in Language Models

Large language models serve speakers of hundreds of languages, yet most of what is known about their inner workings comes from English. Recent interpretability research shows that abstract concepts, including emotions, are encoded inside these models in forms that can be read out and adjusted, changing the tone and behaviour of what a model writes. Whether this holds beyond English is largely unknown.

This project studies emotion concepts across English and Bangla. The central question is whether a model holds one shared concept of fear or joy, or separate versions tied to each language. Work on multilingual representation has focused on word meaning and translation, while work on internal control is almost entirely monolingual. This project addresses the gap between them.

Several open problems follow. One is locating where in a model concepts become independent of language, and where language-specific structure returns. Another is whether an adjustment made in one language carries into another without harming fluency or pushing the output into a different language. A third is whether safety-relevant behaviour responds consistently across languages, where an inconsistency becomes a direct risk to non-English users.

The answers will show whether tools built to monitor and control models in English can be relied on for Bangla and other underrepresented languages.

Team: Md Mubtasim Ahasan, AKM Moshiur Rahman Mazumder, Sami Ibn Rashid, Amin Ahsan Ali