Skip to content
View Vinay-15's full-sized avatar
:octocat:
:octocat:

Highlights

  • Pro

Block or report Vinay-15

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Vinay-15/README.md

header

Typing SVG

👨‍💻 About Me

I'm an AI Infrastructure / LLM Serving engineer and MS in Data Science candidate at CU Boulder (GPA 4.0/4.0). I focus on making LLM inference fast, reliable, and memory-efficient under load — benchmarking and fixing the points where KV cache, throughput, and latency break down on constrained GPUs.

  • 🔭 Building KV-cache-aware LLM serving on vLLM (PagedAttention, preemption, prefix caching)
  • 🛠️ Interested in inference optimization, GPU memory management, and serving-stack reliability
  • 🏥 Ex-Stanford Health Care (AI Dev Intern) · 🏦 Ex-Oracle (2 yrs, core banking at scale)
  • 🏆 Red Bull Innovation Track Winner — HackCU12
Terminal GIF

⚙️ Focus: AI Infrastructure & LLM Serving

  • KV cache & memory management — PagedAttention block pools, preemption/recompute, cache saturation
  • Serving stacks — vLLM, OpenAI-compatible APIs, concurrent load testing, prefix caching, chunked prefill
  • Performance eval — throughput, TTFT, p50/p99 latency, Prometheus /metrics instrumentation
  • GPU & HPC — NVIDIA A100 (MIG), CUDA version matching, Slurm scheduling, Apptainer/Docker containers

vLLM · PagedAttention · CUDA · PyTorch · Slurm/HPC · Apptainer/Docker · Prometheus · FastAPI


🧠 Broader Skills

LLMs & Agents: LangChain · LangGraph · RAG (FAISS + BM25 + RRF) · ReAct Agents · Pydantic · Multi-Agent Systems
ML & Data: PyTorch · TensorFlow · Scikit-learn · Pandas · NumPy · PySpark
Cloud & MLOps: AWS · GCP · Azure · Docker · Podman · Kafka · Airflow · Databricks · Snowflake
Languages: Python · Scala · C/C++ · PL/SQL · Shell/UNIX · SQL


💼 Experience

🏥 AI Developer Intern — Stanford Health Care (May 2026 – Jul 2026 · 3 mo)

  • Built a real-time BACnet/IP data pipeline and database from scratch for hospital telemetry feeding ML/agentic workflows
  • Deployed on-premise LLM agents with tool-calling, async concurrency, and hardened error handling; owned LLM-serving-stack uptime and health monitoring in a regulated environment

🏦 Software Developer — Oracle (Financial Services) (Aug 2023 – Aug 2025 · 2 yrs)

  • Owned core banking modules serving 600+ institutions (UBS, Access Bank); built data pipelines in Python/PL/SQL/Shell/REST
  • Automated an ETL + validation framework (10K+ daily txns, defects −30%); ran containerized deployments (Docker, Podman) at 99.9% uptime

📊 Data Engineer Intern — The Thick Shake Factory (May 2022 – Jul 2022 · 3 mo)

  • Built high-throughput ETL/ELT streaming topologies on Apache Kafka with parallel ingestion via Airflow DAGs, consolidating data from 40+ outlets into columnar storage
  • Optimized cluster utilization to cut pipeline bottlenecks, powering real-time dashboards over 100K+ daily transactions

🤝 Connect with Me

footer

Pinned Loading

  1. Smart-Glove-Sign-Language-Interpreter Smart-Glove-Sign-Language-Interpreter Public

    Developed a smart glove using Arduino and Flex sensors which interprets the sign language through the movement of flex sensors.

    C++ 1 1

  2. KV_Cache_LLM_Serving KV_Cache_LLM_Serving Public

    KV-cache-aware LLM serving on GPU/HPC

    Python

  3. llm-semantic-book-recommender llm-semantic-book-recommender Public

    Forked from t-redactyl/llm-semantic-book-recommender

    The code to accompany the freeCodeCamp tutorial explaining how to use large language models to build a semantic book recommender.

    Jupyter Notebook

  4. Plant-Disease-Detection Plant-Disease-Detection Public

    Jupyter Notebook

  5. Ricoh_GPT_RAG Ricoh_GPT_RAG Public

    Forked from DaSSA-Hackathon-2026/Ricoh_GPT_Data_Dons

    Hybrid RAG Retrieval Pipeline for Ricoh Technical Documentation

    Python

  6. Technical_analyst_Agentic_AI Technical_analyst_Agentic_AI Public

    Python