CV / Resume

Research experience, education, selected publications, and technical expertise. Download the latest resume using the PDF icon.

Professional Summary

  • Research scientist focused on improving the efficiency of LLM reasoning and multimodal models through reasoning-token compression, quantization, and model compression; enhancing VLM grounding and reducing hallucinations through supervised fine-tuning, distillation, and reinforcement-learning-based post-training; and developing calibrated, reliable, and trustworthy VQA systems.
  • Author of 16+ peer-reviewed publications at ICML, NeurIPS, ICLR, CVPR, KDD, AAAI, AISTATS, and other venues.
  • EB-1B Outstanding Researcher (2025).

General Information

Full Name Hitesh Sapkota, Ph.D.
Location Sunnyvale, California, USA
Email hiteshsapkota@gmail.com
Profiles Google Scholar · LinkedIn · GitHub
Languages English, Nepali

Experience

  • Dec 2023 – Present
    Applied Scientist II, Devices & Services — Edge AI & Science
    Amazon, Sunnyvale, California
    • Proposed TL;DR, a reasoning-compression approach using thinking gradients and self-annealing rewards; reduced inference tokens by up to 32% on math and 20% on code while matching or exceeding baseline accuracy. [Under review]
    • Proposed V-STAR, a GRPO-based post-training framework for grounded VLMs; reduced object-level hallucination by 23–44% and sentence-level hallucination by 60–73% across four architectures with no additional inference cost. [Under review]
    • Co-led a study of how token compression and inference-time interventions compose in VLMs, achieving 89% token reduction and 54% hallucination reduction. [Under review]
    • Developed video-aware SigLIP2 encoders that combat embedding collapse in static-camera retrieval, improving text-to-frame Recall@1 by up to 38.2 percentage points. [Under review]
    • Delivered an invited tutorial on Efficient Edge VLMs at AMLC 2025 and shipped a multi-stage edge-VLM training pipeline adopted by multiple teams.
  • 2021 & 2022
    Applied Scientist Intern
    Amazon, Sunnyvale, California & Seattle, Washington
    • Designed attention-based embedding adaptation for sensor technologies, improving performance by 6% over the existing baseline.
    • Built machine-learning models for early detection of AWS service failures, improving performance by 4% over the existing baseline.
  • Aug 2017 – Dec 2023
    Research Assistant
    Machine Learning and Data Intensive Computing Lab, Rochester Institute of Technology
    • Published 10+ papers, including first-author work on distributionally robust optimization, evidential learning, weak supervision, anomaly detection, and network sparsification.
    • Developed methods for calibrated sparse networks, few-shot open-set recognition, and robust weakly supervised anomaly detection.

Education

  • 2017 – 2023
    Ph.D. in Computing and Information Sciences
    Rochester Institute of Technology, Rochester, New York
    • Dissertation: Robust Weakly Supervised Learning for Real-World Anomaly Detection.
  • 2012 – 2015
    B.E. in Electronics and Communication Engineering
    Tribhuvan University, Institute of Engineering, Pulchowk Campus, Lalitpur, Nepal

Selected Publications

Honors and Awards

  • 2025
    • EB-1B Outstanding Researcher — U.S. permanent residency granted for outstanding contributions to machine learning.
  • 2017 – 2023
    • RIT Ph.D. Merit Scholarship — full financial support for doctoral studies.
  • 2022
    • KDD Travel Award.
  • 2014
    • Ncell Scholarship and Excellence Award.

Academic Service

  • Reviewer (2020–2026): NeurIPS, ICML, ICLR, CVPR, ECCV, ICCV, NAACL, EMNLP, KDD, AAAI, IJCAI, MICCAI, Neurocomputing, and IEEE Transactions on Cognitive and Developmental Systems.

Technical Expertise

  • Languages and Frameworks
    • Python, C/C++, Java, PyTorch, Hugging Face Transformers, vLLM, DeepSpeed, TensorFlow, ONNX.
  • LLMs and Multimodal AI
    • LLM reasoning, vision-language models, visual question answering, supervised fine-tuning, distillation, reinforcement-learning-based post-training (GRPO), reasoning-token compression, VLM grounding, and hallucination reduction.
  • Efficient and Reliable ML
    • Quantization and model compression, network pruning and sparsification, uncertainty-aware learning, evidential deep learning, calibration, and distributionally robust optimization.
  • Systems
    • Distributed training, edge and on-device deployment, model-quantization pipelines, and multi-stage VLM training.