Sociolinguistic Infrastructure for Evaluating Real-World AI Behavior
I design language-centered AI evaluation tools and mixed-methods experiments that study how language models express values, stance, politeness, authority, and cultural assumptions in real-world interaction.
My work sits at the intersection of sociolinguistics, discourse analysis, multilingualism, and AI evaluation. I build benchmarks, annotation frameworks, privacy-aware analysis pipelines, and research tools that help teams examine model behavior more carefully across languages, registers, and social contexts.
This portfolio presents four projects that reflect that approach. Together, they show my take on how sociolinguistic insight can become technical infrastructure for studying AI systems in ways that are rigorous, practical, and useful for research, product, and policy.
About
I am a professor of sociolinguistics transitioning from academia into industrial AI research. My academic work has focused on language, style, discourse, identity, branding, and multilingual communication, and more recently on prompt engineering, model behavior, and the sociolinguistics of human-AI interaction. My monograph titled Style as the Steering Wheel: Sociolinguistics Meets Prompt Engineering is under contract with Cambridge University Press. I also have the Responsible Generative AI Specialization from the University of Michigan.
What draws me to AI is not only what models can produce, but how they produce it: what kinds of voice they normalize, what assumptions they treat as neutral, how they adapt across audiences and languages, and where their limits reveal deeper social and epistemic patterns.
I am especially interested in advice, workplace writing, intercultural communication, multilingual prompting, and real-world human-AI use. My goal is to build evaluation systems that make these dynamics visible and actionable.
Featured Projects
A multilingual evaluation suite for values, stance, and sociolinguistic variation in AI advice
Styling Alignment Bench evaluates how language models express values, directness, politeness, hedging, and cultural sensitivity across English, Greek, and Arabic prompts. The benchmark focuses on advice-oriented scenarios such as workplace conflict, education, emotional support, and intercultural misunderstanding.
This project combines benchmark design, structured annotation, automated scoring, and dashboard-based error analysis. It treats alignment not only as a safety question, but also as a sociolinguistic one: how models sound, what kinds of authority they project, and which norms they reproduce when they respond to users.
Mini-Clio for Language and Misuse Patterns
A privacy-preserving clustering tool for discovering sociolinguistic risks in user-AI interaction
Mini-Clio is a lightweight analysis pipeline for identifying patterns in user-AI conversations without depending on raw-text inspection. It combines anonymization, embedding-based clustering, auto-labeling, and interface-based browsing to help researchers detect misuse, register mismatch, culturally loaded hostility, and other language-based risks at scale.
This project is designed as a technical system for pattern discovery. It translates a sociolinguistic research question into a usable research tool.
A mixed-methods study of how professionals use AI for writing, planning, translation, and judgment.?
AI at Work Interviewer examines how people use AI in everyday professional tasks and how that use reshapes expertise, task boundaries, and professional identity. The project combines structured interviews, human coding, and model-assisted thematic analysis to study how users delegate, revise, resist, or rely on AI in their work.
This project reflects my interest in future-of-work questions, but through a language-centered lens: how people use AI not only to save time, but to negotiate voice, confidence, responsibility, and self-presentation.
A privacy-preserving pipeline for structural efficiency and sustainability in user–AI interaction
Prompt Circularity identifies and analyzes recurring inefficiencies in prompt design, such as redundant specifications, repair loops, and multilingual token inflation, by abstracting interaction patterns into reusable structures. The system operates without exposing raw text, focusing instead on discourse-level representations.
This project combines structural abstraction, clustering, inefficiency detection, and optimization reporting. It reframes sustainability as a sociolinguistic issue, examining how prompting practices shape computational cost, interaction quality, and the environmental footprint of real-world AI use.
What I Build
I build technical and research artifacts that help people inspect model behavior more clearly:
multilingual benchmarks
annotation frameworks
privacy-aware analysis pipelines
mixed-methods evaluation studies
dashboards for qualitative and quantitative error analysis
research reports that connect findings to product and policy implications
Research Questions That Guide My Work
How do language models express values in ordinary interaction?
What kinds of voice, authority, and politeness do they normalize by default?
How do these patterns shift across languages, audiences, and discourse settings?
Where do models flatten cultural difference or treat one style as neutral?
How can evaluation tools make these patterns visible in ways that support better research, safer products, and more socially aware design?
Methods and Tools
My work combines qualitative and computational approaches, including discourse analysis, sociolinguistic coding, multilingual prompt design, corpus analysis, mixed-methods interviewing, embeddings-based clustering, and lightweight interface development.
I work with Python-based research pipelines and build projects that are meant to be inspectable, reproducible, and useful to both technical and non-technical collaborators.
Why This Portfolio Exists
Many AI evaluations focus on whether a system is accurate, safe, or efficient. Those questions matter. But they do not fully explain how a model behaves socially, what kinds of norms it reproduces, or why certain responses feel supportive, distancing, overconfident, flattening, or culturally tone-deaf.
This portfolio explores those questions by treating AI output as both a technical object and a social one.
Contact
I am interested in research roles focused on societal impacts of AI, human-AI interaction, evaluation, multilingual AI, and language-centered model behavior.
Personal website: Irene Theodoropoulou
LinkedIn: https://www.linkedin.com/in/irene-theodoropoulou-a12aa611/
Email: theodoropouloui21@gmail.com