✉️ anuragbeniwal09@gmail.com
I am a Member of Technical Staff at ElevenLabs, where I lead the development of ElevenAgents for Support, the company's enterprise customer-support voice and chat agent platform.
Prior to ElevenLabs, I led AI for worldwide customer service at Amazon, where I drove the transition from hand-engineered workflows to enterprise agents handling approximately 2 billion customer interactions per year across 400 million unique customers in North America, Europe, and India. Embedded AI into human associate workflows, extending beyond tooling to CX, training, and change management across more than 100,000 support associates. Earlier in my time at Amazon and AWS, I was a founding member of several products that are now core to Amazon's AI portfolio, including AWS SageMaker, Amazon Go/Style, Prime Wardrobe.
I believe in a top-down approach to research, motivated by challenges and insights observed at scale in industrial deployment, and using research as a tool to solve those problems rather than the other way around. I have led both small, nimble teams and broader cross-functional organizations of roughly 50 applied researchers and engineers across applied-research and engineering functions at Amazon and ElevenLabs, and am passionate about building teams, developing people, and amplifying impact while remaining personally engaged with the substance of the work.
My early career focused on applied research in recommender systems and tensor factorization in production settings, including large-scale tensor factorization, personalized compatibility in fashion via subspace attention, and Outfittransformer, one of the first papers to model outfits holistically as a single representation using transformers.
More recently, I have worked on task-oriented dialogue systems and large language models. My contributions include neural machine translation for low-resource languages, self-supervised methods for dialogue summarization, a study showing that explicit chain-of-thought reasoning can significantly degrade instruction-following accuracy in goal-oriented tasks together with a constraint attention metric to quantify the degradation (NeurIPS 2025 Spotlight), a principled measure of internal and external uncertainty in LLMs to decide when an agent should ask a clarifying follow-up (AAAI 2026), and benchmarks for evaluating policy- and instruction-following in multi-turn task-oriented dialogue. Much of this work was motivated by my hands-on experience post-training and deploying customer-support agents since early 2023.