- KV cache & memory management — PagedAttention block pools, preemption/recompute, cache saturation
- Serving stacks — vLLM, OpenAI-compatible APIs, concurrent load testing, prefix caching, chunked prefill
- Performance eval — throughput, TTFT, p50/p99 latency, Prometheus
/metricsinstrumentation - GPU & HPC — NVIDIA A100 (MIG), CUDA version matching, Slurm scheduling, Apptainer/Docker containers
vLLM · PagedAttention · CUDA · PyTorch · Slurm/HPC · Apptainer/Docker · Prometheus · FastAPI
LLMs & Agents: LangChain · LangGraph · RAG (FAISS + BM25 + RRF) · ReAct Agents · Pydantic · Multi-Agent Systems
ML & Data: PyTorch · TensorFlow · Scikit-learn · Pandas · NumPy · PySpark
Cloud & MLOps: AWS · GCP · Azure · Docker · Podman · Kafka · Airflow · Databricks · Snowflake
Languages: Python · Scala · C/C++ · PL/SQL · Shell/UNIX · SQL
🏥 AI Developer Intern — Stanford Health Care (May 2026 – Jul 2026 · 3 mo)
- Built a real-time BACnet/IP data pipeline and database from scratch for hospital telemetry feeding ML/agentic workflows
- Deployed on-premise LLM agents with tool-calling, async concurrency, and hardened error handling; owned LLM-serving-stack uptime and health monitoring in a regulated environment
🏦 Software Developer — Oracle (Financial Services) (Aug 2023 – Aug 2025 · 2 yrs)
- Owned core banking modules serving 600+ institutions (UBS, Access Bank); built data pipelines in Python/PL/SQL/Shell/REST
- Automated an ETL + validation framework (10K+ daily txns, defects −30%); ran containerized deployments (Docker, Podman) at 99.9% uptime
📊 Data Engineer Intern — The Thick Shake Factory (May 2022 – Jul 2022 · 3 mo)
- Built high-throughput ETL/ELT streaming topologies on Apache Kafka with parallel ingestion via Airflow DAGs, consolidating data from 40+ outlets into columnar storage
- Optimized cluster utilization to cut pipeline bottlenecks, powering real-time dashboards over 100K+ daily transactions
