Production ML · 2026
Offline translation in Microsoft Edge
Quantized and integrated translation models for private, on-device translation through the built-in Translator API.
Research Scientist at Microsoft · ML Systems
I build AI systems that survive the trip from research to production.
My work spans multilingual and multimodal models, data and distributed training, evaluation, quantization, GPU kernels, and low-latency inference.
tgowdan [at] gmail [dot] comFeatured work
Selected outcomes across production AI, model inference, algorithms, and systems software.
Production ML · 2026
Quantized and integrated translation models for private, on-device translation through the built-in Translator API.

Research leadership & systems · 2025–2026
I lead the WMT Model Compression shared task and built its two-year evaluation infrastructure for runnable systems, controlled H100 execution, and quality–footprint–speed trade-offs. TahomaMT tests those ideas in a custom C++/CUDA runtime.
Systems software · 2026
A thread-safe C++23 compression core with multi-language bindings, native ZIP, and a fast PNG codec.
Algorithms · 2026
Replaced repeated linear scans with a heap-based update and merged the implementation upstream into Google SentencePiece.
Earlier systems
Research infrastructure that became useful beyond the work that started it.
Multilingual data and training
I built an open toolchain for reproducible MT data, vocabularies, training, and inference, then used it to train a single model spanning more than 500 source languages. As a WMT General MT organizer from 2022 through 2026, I have maintained versioned MTData recipes that give participants reproducible constrained-track datasets across five editions.
Open model deployment
A practical web interface, REST API, and batch decoder for deploying Meta's No Language Left Behind models across 200 languages. It remains my most widely adopted personal repository.
Distributed systems
I created Sparkler at USC: an extensible web crawler built around Apache Spark, Kafka, Solr/Lucene, Tika, and distributed JavaScript rendering. I later handed the project to its maintainers when I shifted focus to my Ph.D.
Research
Proceedings of the AAAI Conference on Artificial Intelligence
Notes
Quality accounting, W4A8 kernels, and one-H100 inference results.
One accelerated core across data pipelines, Docker, PNG, ZIP, and the browser.
The algorithm, implementation, benchmarks, and upstream journey.
Background
At Microsoft, I work on multilingual and multimodal AI and the systems that train and serve these models. I have also helped organize the WMT General MT shared task for five consecutive editions, from 2022 through 2026. Previously, I spent five years as a Research Engineer at USC ISI and worked as a Data Scientist intern at NASA JPL.
I earned my Ph.D. and M.S. in Computer Science from USC. My doctoral work studied rare phenomena learning in neural machine translation (dissertation). I have also built software at startups, co-founded Datoin, and served as an Apache Software Foundation committer and PMC member.
Invited talks
Delivered Efficient Software for AI Systems: From Training to Inference at the 4th IEEE International Conference on Knowledge Engineering and Communication Systems, hosted by SJC Institute of Technology on April 24, 2026. Slides Program
Presented LLMs and Multilingual AI: Enabling Vernacular Innovations in India for faculty and graduate students in SJCIT's Faculty Development Program on January 22, 2026. Slides
Beyond the work
I grew up in a farming family in southern India, where work was measured in acres and yield, not hours. That upbringing still shapes how I approach hard problems: understand the whole system, stay practical, and judge the work by what it produces.