Skip to content
View vishalhabib99's full-sized avatar
🎯
AI PM — I ship agent evals, then red-team them in public
🎯
AI PM — I ship agent evals, then red-team them in public

Block or report vishalhabib99

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
vishalhabib99/README.md

Hi, I'm Vishal 👋

Lead AI Product Manager & Builder — Agentic AI · leading the 0→1 Agentic AI Digital Advisor at Vanguard · ex-T-Mobile, eBay · Stanford GSB

10+ years launching 0→1 products and scaling platforms 1→100 across fintech, SaaS and enterprise platforms, and marketplaces: $9B+ in revenue platforms, 29M+ MAU, $300M+ in cost savings. On the side I build the eval and trust tooling I wish every AI team had, and I publish the failures along with the passes.

▶️ 76-second walkthrough (AI-narrated):

vishal_walkthrough_titled.mp4

What ties it together: AI output is only worth shipping when it can be checked against a source of truth, whether that's an IRS rule, a tool's own schema, or a platform's API contract.

⏱️ Got 2 minutes? Start here:

  1. A real report from a real user: an outside maintainer opened an issue and a bot graded their MCP server. No install needed.
  2. A red team that broke my own checker: all 22 in-scope attacks got through, and I published it as-is.
  3. Try a checker yourself: an agent-to-agent handoff, decided ACT, ESCALATE or BLOCK, in your browser.
  4. A 155K★ project built its fix on my PR: Langflow's maintainer adapted my fix for MCP tool arguments and credited me as co-author.

✅ Proof from outside

None of these maintainers work with me.

"Thank you for the exact tools/list results and for explaining the effect on clients that use annotations for approval." (codebase-memory-mcp maintainer)
"Thanks for shipping this and keeping the regression case." (an outside tester whose agent found the bug)

🗺️ Where my work sits in the AI stack: what I shipped at work vs. built in the open, layer by layer
Layer Shipped at work Built in the open
Apps & human-in-the-loop 🏦 Vanguard Digital Advisor experience · 🏢 T-Mobile agentic support platform 🏦 Contribution room calculator · 🏦 Human review queue · 🛒 Listing checker demo
Agents & orchestration 🏦 Agentic AI Digital Advisor (0→1) · 🏢 Autonomous enterprise agent platform GuardedSession: ACT, ESCALATE or BLOCK on each live tool call · 🏢 agent-handoff-check: authority can only narrow across handoffs
Tools & APIs (MCP) 🛒 eBay API standardization across hundreds of teams mcp-doctor, mcp-fuzz, mcp-reality-check · 🏦 check_answer and contribution_room MCP tools · 🛒 check_listing MCP tool
Evals & quality 🏦 Model evals for correctness, groundedness, safety, latency · 🏢 IntentCX evaluation framework 🏦 Blind, pre-registered evals · 🛒 Failed blind run, fresh pass, then broken by a red team · 🏢 RAG prototype on blind tickets: accuracy fell from 60% to 38% · ⚛️ Quantum circuit grader: right on all 90 real answers, yet 4 red teams got 27 wrong circuits passed (write-up) · /eval-plan
Guardrails, governance & risk 🏦 FINRA/SEC-compliant responsible AI design · 🏢 Governance aligned to NIST AI RMF 🏦 Model risk pack (SR 26-2 + NIST AI 600-1) · 🏦 Prompt-injection red team, before and after the fix · mcp-trust-check release gate · 🏢 Red-teamed handoff checks
Cost & pricing 🏢 Accuracy, cost and latency tuned per interaction type 🏢 Which model, and how to price it
Product decisions 0→1 strategy, launch gates, adoption and containment metrics PRD: Agent Outcome Trust Score · /build-or-not · Agent Readiness Scorecard · What I decided not to build · Who owns trust for agent tools

🧱 Portfolio

🏦 Fintech

  • 🛡️ retirement-answer-check: checks an AI's draft answer to a retirement-account question before a customer sees it, and decides SEND or REVIEW with an IRS or FINRA source for every flag.
    Pattern rules alone let 5 of 15 blind wrong facts through; adding a fact-checking judge brought that to 0 of 25. Also: contribution room calculator · model risk pack

🛒 Marketplaces

  • 🏷️ listing-claim-check: checks an AI-written listing against the seller's own item specifics and decides PUBLISH or REVIEW. Try it →
    v0.1 failed its blind run and v0.2 passed a fresh one, then a red team got all 22 attacks through. Published as-is.
  • 🧩 eBay case study: the platform behind $300M+ in savings, and why clear contracts matter for agents.

🏢 SaaS & enterprise platforms

  • 🔗 agent-handoff-check: checks every agent-to-agent handoff against what the customer authorized and decides ACT, ESCALATE or BLOCK. Try it →
    A red team got 3 unauthorized calls through; after one design change, a fresh blind run let 0 of 18 through and blocked 0 of 14 legitimate calls.
  • 📡 T-Mobile case study: four architecture decisions behind the agentic platform (25M users, 60% containment) and what I'd do differently.
  • 📊 Which model, and how to price it: model cost is under 5% of the value delivered, so the real constraint is draft quality.

🧰 Trust tooling for the tools AI agents call (MCP)

mcp-doctor scanning homeassistant-ai/ha-mcp: Quality 96% grade A, Security 98% grade A, 88 tools found, with per-tool OK and WARN lines

🧭 For AI product managers

✍️ Writing

More writing

💼 Career, by vertical

Vanguard · T-Mobile · eBay · earlier roles
  • 🏦 Fintech: Vanguard. Product strategy for Digital Advisor and Personal Advisor ($6B+ LOB, 4M+ MAU) at the world's second-largest asset manager (~$12T AUM). Architecting the Agentic AI Digital Advisor from 0→1: LLM orchestration, agent workflows, model evals and FINRA/SEC-compliant responsible AI. Also launched T-Mobile Money and, earlier, digital lending modernization at Axis Bank.
  • 🏢 SaaS & enterprise platforms: T-Mobile. Product lead for T-Life, the flagship app ($3B+ LOB, 25M+ MAU); managed 3 PMs and a 40+ person org. Launched one of the first autonomous enterprise Agentic AI platforms in US telecom: 75% adoption, 60% containment, 80% CSAT, 30% fewer support calls. Built the IntentCX AI governance and eval framework (NIST AI RMF), adopted by 3 more teams. AI personalization: +27% engagement, +15% conversion.
  • 🛒 Marketplaces: eBay. Drove $300M+ in savings and +35% adoption through platform modernization and API standardization across hundreds of teams; 40% faster deploys, 99.9% availability on seller APIs.
  • 🩺 Earlier, healthcare: Premera Blue Cross. Billing and payment redesign that cut task completion time 25%.
Recognition
  • Patent filed: sole inventor on a U.S. provisional patent application (No. 63/980,243, May 2026) for agentic AI orchestration and governance.
  • Top Product Leader, T-Mobile (2025), for building and launching the enterprise agentic AI platform.
  • Keynote speaker, T-Mobile Technology Innovation Summit (2025): AI strategy keynote to 1,200+ attendees.
Education & certifications
  • Stanford Graduate School of Business: Executive Program, Harnessing AI for Breakthrough Innovation and Strategic Impact (2026)
  • MBA, Business Strategy and Marketing, Indiana University of Pennsylvania · MS, Information Technology Management, Campbellsville University
  • Certifications: PMP · SAFe POPM · AWS Solutions Architect Associate · CSPO · PMI-PBA · Google AI Essentials
What this portfolio doesn't cover
  • Semantic hallucination ("is this answer true" in general): it needs an LLM judge, which would break the deterministic, no-API-cost design of the MCP tools. retirement-answer-check uses a judge only for narrow fact checks against IRS and FINRA sources.
  • Live agent red-teaming: mcp-doctor's security score audits a tool's code and description for injection risk, not whether a live agent can be manipulated at runtime.

📫 Reach me

Email · LinkedIn · Website · dev.to

Pinned Loading

  1. mcp-doctor mcp-doctor Public

    Static audit of MCP servers for what breaks an agent calling them. Part of a 3-tool QA suite: mcp-trust-check runs all three. Catches missing descriptions, undocumented params, unhandled errors, co…

    Python 3

  2. retirement-answer-check retirement-answer-check Public

    Checks AI-drafted answers to US retirement-account questions (SEND or REVIEW, with an IRS or FINRA source for every flag), plus a 2026 contribution-room calculator and MCP tool.

    Python 1

  3. listing-claim-check listing-claim-check Public

    Checks AI-written marketplace listings against the seller's item specifics before they're published: PUBLISH or REVIEW, with the exact words behind every unbacked claim. Deterministic, blind-evalua…

    Python 1

  4. agent-handoff-check agent-handoff-check Public

    Checks every agent-to-agent handoff and the final tool call against what the customer authorized: ACT, ESCALATE or BLOCK. Authority can only narrow. Deterministic, blind-evaluated, red-teamed.

    Python 1

  5. ai-pm-skills ai-pm-skills Public

    Claude Code skills for AI product managers: /build-or-not, /eval-plan, /agent-trust-review. Each tested with evals, failures included.

    Shell 1

  6. circuit-claim-check circuit-claim-check Public

    Checks whether a quantum circuit an AI agent wrote is the circuit it was asked for: exact Qiskit grading, no model in the loop. Includes four red-team rounds and every finding.

    Python 1