Senior AI/ML Engineer
Labelbox
Translating frontier-lab requirements into production RL data and environments; evaluating capability gaps across terminal, software, and cognitive tasks.
Vikram Kharvi · Applied AI researcher + engineer
From reinforcement-learning environments to foundation-model training and evaluation—six years turning ambitious research into infrastructure that works.
Currently building RL environments that help LLMs learn from real tasks, feedback, and failure. →01 / Research paths
Three connected bodies of work: how models learn, how they improve through feedback, and how they run reliably in production.
Foundations
Data, tokenization, model architecture, and the systems that create foundation models.
Coming soon ↗ 02TrainRL
Reinforcement learning, environments, graders, evaluations, and specialist agents.
Read the public post-training blog ↗ 03Serving
Paged KV caching, continuous batching, fused kernels, and the systems work behind reliable low-latency serving.
Coming soon ↗02 / Experience
Research, data, infrastructure, deployment—the hard problems rarely respect team boundaries.
Labelbox
Translating frontier-lab requirements into production RL data and environments; evaluating capability gaps across terminal, software, and cognitive tasks.
HTC Inc
Built hot-swap model adapters that cut deployment time by 80% and architected an end-to-end AI workflow management system.
Usable Machines
Built on the WhiteRabbitNeo foundation model and reduced task perplexity from 13 to 5 through preference optimization and task-specific fine-tuning.
Northwestern University
Created a build-an-LLM course and researched controlled legal-text generation with T5, reaching a ROUGE-L F1 score of 0.44.
Code Techniq
Led an eight-person engineering team, built browser-based vision systems, and designed NLP pipelines for knowledge and medical applications.
03 / Point of view
I care about the full learning loop: the data we choose, the worlds we ask models to navigate, the failures we measure, and the systems that carry those lessons into production.
Research should leave the notebook.
— VK
04 / Ideas from the workbench
Notes on alignment, RL environments, and the signals that help language models improve through experience.
A map of supervised tuning, preferences, reward models, RL, safety, agent trajectories, distillation, and continual improvement.
Read the current blog ↗Current research focus
The full pre-training, post-training, and inference libraries are preserved. For now, only the foundational post-training guide is public; future field notes will appear one at a time through the research feed.
05 / About
I’m Vikram Kharvi, an applied AI researcher and engineer in California. I’ve worked across the AI lifecycle—from training and preference optimization to evaluation, deployment, and monitoring.
I hold a Master of Artificial Intelligence from Northwestern University and a Bachelor of Information Science Engineering from PES University.
06 / Research feed
Subscribe for new field notes on reinforcement learning, LLM post-training, agent evaluation, and efficient inference systems.
Subscribe via RSS Works with Feedly, Inoreader, NetNewsWire, and other feed readers.