AI Foundations team, Cambridge, Massachusetts, USA
Research on efficient and production-ready Large Language Model training, post-training and inference
Red Hat Inc.
Senior Machine Learning Research Engineer
Oct 2025 – Sep 2026
Drove research on post-training and inference, including mixed-precision quantization, parallel token drafting for speculative decoding, RL rollouts, and Agentic AI evaluation (e.g., SWE-Bench)
Contributed to open-source frameworks: vLLM, llm-compressor, speculators and the Red Hat AI Hugging Face model repository, enabling production deployment of compressed and accelerated LLMs
Integrated FP8 quantized Inkling, Qwen 3.6, Granite4 and Nemotron family of models into vLLM and llm-compressor; developed Dflash drafters (parallel token predictors) for Qwen3 and Nemotron
Worked on tool calling and Agentic datasets and evaluation
Mentored one intern on the KV-Dflash project
Argonne National Laboratory
Postdoctoral Researcher
Aug 2023 – Sep 2025
AI/ML team of Argonne's Leadership Computing Facility (ALCF)
Research at the intersection of Systems and Deep Learning; optimizing inference and finetuning of LLMs
Collaborated with scientists and engineers from Nvidia, Intel, AMD, SambaNova, Cerebras, Groq, Graphcore and Habana
Argonne National Laboratory
Research Intern
Sep 2021 – Nov 2021
AI/ML team of Argonne's Supercomputing facility
Worked on a project at the intersection of Pruning, Quantization and Neural Architecture Search