Skip to content
hitensjPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Kalidas — AI Creative Co-Director

Kalidas is named after the legendary Sanskrit poet and dramatist who flourished during the Gupta Dynasty — often called the golden age of Indian art, literature, and science. Kalidasa is remembered as one of the greatest wordsmiths in history, celebrated for works like Abhijnanashakuntalam, Meghaduta, and Raghuvamsha, where he turned simple briefs — a season, a longing, a myth — into masterpieces of imagery, rhythm, and emotional precision.

This project channels that same spirit for the modern creator: a multi-agent AI creative co-director built for the IBM AI Builders Challenge (Creative Industries track) that takes a raw brief and, like a co-writer with an eye for hooks, visuals, and rhythm, shapes it into a fully polished content package — hooks, script, visual prompts, audio suggestions, and a storyboard — in a single automated pipeline. A second pipeline runs on a daily schedule to deliver a Morning Scroll digest email packed with trending hashtags, visual themes, and concepts sourced via Gemini's native Google Search grounding.

The system is designed for a single trusted user (you) and restricts access at every layer: Google OAuth at the frontend, ID-token verification in the backend, and a dedicated Cloud Scheduler service account for the internal digest endpoint. All credentials are stored in Google Cloud Secret Manager — no secrets live in code or container images.


Demo & Screenshots

Project Demo Video

Watch a full walkthrough of the Kalidas platform: Kalidas Demo Video

Application Screenshots

Google Sign-In (OAuth Login)

Access restricted to authorized accounts only: Google Sign-In Screen

Workshop Page (New Content Brief Submission)

Input the brief parameters to trigger the multi-agent creation pipeline: Workshop Screen

Architecture

Two LangGraph Graphs

Digest Graph (backend/graphs/digest_graph.py) runs on a Cloud Scheduler cron: START → trend_scout → assemble_digest → send_digest → END

The trend_scout agent calls Gemini 2.5 Flash with Google Search grounding to collect hashtags, visual themes, and concepts. assemble_digest renders the results into an HTML + plain-text email via Jinja2 templates. send_digest dispatches it through the Gmail API.

Creation Graph (backend/graphs/creation_graph.py) runs on demand: START → parse_brief → [hook_smith ‖ prompt_smith ‖ audio_curator] → gather → directors_cut → (revise | export) → END

Three specialist agents run in parallel via LangGraph's Send API. The Director's Cut agent (Gemini 2.5 Pro) critiques the assembled package and either approves it or routes only the weak agents back for targeted revision (up to MAX_REVISION_PASSES times). The export node writes the final package to a Google Doc via the Drive API and returns the Doc URL.

Five Agents

Agent Model Role
Trend Scout Gemini 3.5 Flash Grounded trend research via Google Search
Hook Smith Gemini 3.5 Flash Hooks, platform script, caption copy
Prompt Smith Gemini 3.5 Flash Veo video prompts, Lyria audio prompts, storyboard
Audio Curator Gemini 3.5 Flash Ranked YouTube audio suggestions with rationale
Director's Cut Gemini 3.5 Pro Critique, approval gate, revision routing

MCP Server

mcp_server/ is a stdio-transport MCP server exposing four tools to the LangGraph agents:

  • search_grounded_trends — Gemini + Google Search grounding (API key)
  • youtube_audio_search — YouTube Data API v3 (API key)
  • gmail_send_digest — Gmail API (OAuth2 user credentials)
  • docs_export_package — Drive + Docs API (OAuth2 user credentials)

Frontend

frontend/ is a React/Vite SPA served from Cloud Run via nginx. It offers three views:

  • Login — Google OAuth sign-in (ID token stored in React context only)
  • Workshop — Brief submission form + live progress timeline polling the job status endpoint
  • Archive — Read-only list of recent Morning Scroll digest metadata

Prerequisites

  • gcloud CLI authenticated with gcloud auth login
  • Docker Desktop (or Docker Engine)
  • Node.js 20+
  • Python 3.12+
  • A Google Cloud project with a billing account attached

One-Time Setup

1. Enable APIs and create IAM resources

export GCP_PROJECT_ID=your-project-id
export REGION=us-central1
bash infra/iam_setup.sh

This enables the required Google Cloud APIs, creates the kalidas-scheduler service account, creates all Secret Manager secrets with placeholder values, and grants the backend Cloud Run runtime identity access to those secrets.

2. Fill in real secret values

Replace each placeholder secret created by iam_setup.sh:

# Example — repeat for every secret listed in infra/iam_setup.sh
echo -n "AIza..." | gcloud secrets versions add gemini-api-key --data-file=-

Secrets to populate:

Secret name Description
gemini-api-key Gemini API key from Google AI Studio
youtube-api-key YouTube Data API v3 key
gmail-sender-address Gmail address that sends digest emails
digest-recipient-email Email address that receives digests
google-client-id OAuth2 Client ID (Web application type)
google-client-secret OAuth2 Client Secret
allowed-google-account The single Google account email allowed to use the Workshop
google-drive-folder-id Google Drive folder ID for content package exports
google-oauth-token-json OAuth2 token JSON for Gmail + Drive (see step 3)
max-revision-passes Integer, default 2
frontend-url Set automatically by infra/deploy.sh after first deploy

3. OAuth2 consent flow for Gmail + Drive

Gmail and Drive operations require personal OAuth2 user credentials (service accounts cannot send Gmail as a personal address or write to a personal Drive without Workspace domain delegation).

# 1. In Google Cloud Console → APIs & Services → Credentials:
#    Create an OAuth 2.0 Client ID of type "Desktop app". Download client_secret.json.

# 2. Run the one-time consent flow (opens browser):
python mcp_server/scripts/oauth_consent.py --client-secrets client_secret.json

# 3. Store the resulting token.json in Secret Manager:
gcloud secrets versions add google-oauth-token-json \
  --data-file=token.json \
  --project="${GCP_PROJECT_ID}"

Local Development

1. Copy the example env file

cp .env.example .env
# Edit .env and fill in all values

2. Start backend + frontend with Docker Compose

docker-compose up --build
  • Backend available at http://localhost:8000
  • Frontend available at http://localhost:5173
  • Vite proxies /api and /internal to the backend automatically

3. Run the MCP server manually (optional)

cd mcp_server
python -m mcp_server.server

Deployment

Build, push, and deploy everything

export GCP_PROJECT_ID=your-project-id
export REGION=us-central1
export VITE_GOOGLE_CLIENT_ID=your-oauth-client-id

bash infra/deploy.sh

The script:

  1. Builds and pushes the backend Docker image to Artifact Registry
  2. Builds and pushes the frontend Docker image to Artifact Registry
  3. Substitutes image references into infra/cloudrun_backend.yaml and deploys
  4. Substitutes image references into infra/cloudrun_frontend.yaml and deploys
  5. Calls infra/create_scheduler.sh to create/update the Cloud Scheduler job
  6. Prints both Cloud Run URLs on completion

After first deployment

Re-run iam_setup.sh once the backend Cloud Run service exists, so the roles/run.invoker binding on the service can be applied:

bash infra/iam_setup.sh

API Reference

Method Path Auth Description
POST /api/create Bearer ID token Submit a ContentBrief; returns {"job_id": "..."} immediately
GET /api/create/{job_id}/status Bearer ID token Poll job progress: {status, current_node, partial_results, final_package}
GET /api/digests Bearer ID token Returns last 10 Morning Scroll digest metadata records
POST /internal/run-digest OIDC (Scheduler SA) Triggers the Digest Graph; called by Cloud Scheduler
GET /health None Health check — returns {"status": "ok"}

ContentBrief schema

{
  "title": "string",
  "platform": "tiktok | reels | youtube_shorts",
  "mood": "string",
  "topic": "string",
  "target_audience": "string",
  "duration_seconds": 30
}

Google APIs Used

API Credential Type Purpose
Gemini API (gemini-2.5-flash / pro) API Key Agent reasoning, Google Search grounding
YouTube Data API v3 API Key Audio/music track search
Gmail API OAuth2 user-consent (refresh token) Sending Morning Scroll digest emails
Google Drive API OAuth2 user-consent (refresh token) Creating export folders for content packages
Google Docs API OAuth2 user-consent (refresh token) Writing content packages as Google Docs
Google Identity (OAuth2) OAuth2 Client ID Frontend sign-in, backend ID-token verification
Cloud Run OIDC service account Cloud Scheduler → backend invocation

Project Structure

kalidas/
├── backend/                  # FastAPI app + LangGraph graphs + agents
│   ├── agents/               # hook_smith, prompt_smith, audio_curator, directors_cut, trend_scout
│   ├── graphs/               # digest_graph.py, creation_graph.py, state.py
│   ├── jobs/                 # job_store.py (in-memory async job tracking)
│   ├── middleware/           # auth.py (Google ID-token verification)
│   ├── routers/              # api.py, internal.py
│   ├── templates/            # Jinja2 email templates
│   ├── config.py             # Centralised pydantic-settings config
│   ├── main.py               # FastAPI app entrypoint
│   ├── requirements.txt
│   └── Dockerfile
├── frontend/                 # React/Vite SPA
│   ├── src/
│   │   ├── pages/            # Login.tsx, Workshop.tsx, Archive.tsx
│   │   └── components/       # ProgressTimeline.tsx, PackageView.tsx
│   ├── nginx.conf
│   ├── package.json
│   └── Dockerfile
├── mcp_server/               # stdio MCP server (four tools)
│   ├── tools/                # trends.py, youtube.py, gmail.py, docs.py
│   ├── scripts/              # oauth_consent.py (one-time consent flow)
│   ├── auth.py               # Credential helpers
│   └── server.py
├── infra/                    # Cloud infrastructure scripts + manifests
│   ├── iam_setup.sh          # One-time IAM + Secret Manager setup
│   ├── cloudrun_backend.yaml # Cloud Run service spec (backend)
│   ├── cloudrun_frontend.yaml# Cloud Run service spec (frontend)
│   ├── create_scheduler.sh   # Creates/updates Cloud Scheduler job
│   └── deploy.sh             # Full build + deploy pipeline
├── docker-compose.yml        # Local development
├── .env.example              # Environment variable reference
└── README.md

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages