My current interests fall into two closely related areas.
Machine learning & LLMs
How can large language, vision, and multimodal models be trained and served efficiently, and how should inference, training, and reinforcement-learning pipelines be built so they scale without losing reliability?
LLMs / computer vision / multimodal / reinforcement learning / inference / training
Systems for machine learning
How do distributed systems, Kubernetes-native infrastructure, and GPU systems combine into a platform that ML workloads can depend on at scale, and how is that platform kept correct as it grows?
Distributed systems / Kubernetes-native ML infrastructure / GPU systems / parallel computing / ML systems
Most of this work is public: the CV lists the projects, pull requests, and talks behind each area.