← flower

research

I believe that as LLMs become increasingly integrated into everyday life, it is important to better understand their social awareness and moral reasoning. Therefore, my research interests are focused on social intelligence, fairness, moral judgment, and behavioral patterns of LLMs in high-stakes social contexts, especially in education. I am interested both in analyzing and improving these capabilities of existing models, as well as in developing LLM-based frameworks that support humans in such domains. I am also interested in these topics in the context of AI Safety and Alignment Research.

publications

preprint

Complexity is complex: Metric-driven iterative text simplification via LLMs

Mariia Eremeeva, Donya Rooein, Sankalan Pal Chowdhury, Mrinmaya Sachan

Abstract: This project focuses on a holistic, metric-based view of complexity with the aim to enable effective LLM-based complexity adjustments.

EACL 2026

PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors

Donya Rooein*, Sankalan Pal Chowdhury*, Mariia Eremeeva, Yuan Qin, Debora Nozza, Mrinmaya Sachan, Dirk Hovy

Abstract: PATS is a framework for developing personality-aware teaching strategies with large language model tutors. The work explores how different personality traits can be leveraged to create more effective and personalized educational experiences.

preprint

Tokenization Matters: Improving Low-Resource Language Modeling for Yakut with Custom Tokenizer

Mariia Eremeeva*, Abu Bakr Rahman Shaik*, Rada Kamysheva*, Nishant Kumar Singh*

Abstract: In this work, we address these limitations for Yakut (Sakha), a Turkic language with Cyrillic orthography, by engineering a Byte Pair Encoding (BPE) tokenizer with Yakut-specific pre- and postprocessing rules and tailored special tokens.

preprint

BERTweetConvFusionNet: Enhancing BERTweet for Twitter Text Sentiment Classification

Debeshee Das*, Mariia Eremeeva*, Piyushi Goyal*, Laura Schulz*

Abstract: This paper proposes a novel deep-learning based solution for Twitter sentiment analysis that addresses the challenges of automatic and noisily labelled data. Leveraging the pre-trained BERTweet model for embeddings, we develop a novel CRNN-based ‘fusion net’ architecture combining CNN, RNN, and Attention layers.

* denotes equal contribution.