Skip to content

Add disaster response coordination agent (Agent Kernel + Gemini) - #628

Open
Jenitha23 wants to merge 4 commits into
yaalalabs:developfrom
Jenitha23:develop
Open

Jenitha23 wants to merge 4 commits into
yaalalabs:developfrom
Jenitha23:develop

Conversation

@Jenitha23

@Jenitha23 Jenitha23 commented Aug 17, 2026 •

Copy link
Copy Markdown

Description

Adds a Gemini-powered submission for the IDEALIZE 2026 Agent Kernel mini-competition: a
three-agent Disaster Response & Resource Coordination system (intake → priority/matching →
dedup/dispatch) that parses free-form need/offer messages, scores urgency, matches across
regions with distance/transport awareness, deduplicates, and dispatches a WhatsApp notification.

Type of Change

  • New feature (non-breaking change which adds functionality)
  • Documentation update
  • Test update

Related Issues

Fixes #
Relates to #

Changes Made

  • Added use-cases/disaster-response-agent/: a 3-agent Agent Kernel pipeline
    (intake_agent → priority_matching_agent → dedup_dispatch_agent) built on the OpenAI
    Agents SDK integration, using Gemini (gemini-3.1-flash-lite) via LiteLLM's native
    integration for all LLM calls
  • Implemented cross-region resource matching weighted by quantity coverage, proximity/distance,
    and transport compatibility ("no transport" / "can deliver" signals detected from message text)
  • Added duplicate detection/merging and a (dummy-mode by default, real-mode configurable)
    WhatsApp dispatch notification, with automatic retry/backoff on rate-limited LLM calls
  • Added demo.py (canonical CLI entry point) and api.py (REST API entry point)
  • Added tests/test_tool_layer.py (32 deterministic unit tests) and tests/test_agent_e2e.py
    (live conversational test via agentkernel.test.Test, skipped automatically without a
    GEMINI_API_KEY), plus test-config.yaml
  • Added SPEC.md and README.md covering the problem, architecture, agent/tool
    responsibilities, memory design, SDG alignment, setup (including Windows-specific notes),
    and limitations

Testing

  • Unit tests pass locally
  • Integration tests pass locally
  • Manual testing completed
  • New tests added for changes

Ran uv run pytest locally: 32 passed, 3 skipped (the live end-to-end tests, which require a
GEMINI_API_KEY). Manually tested the full pipeline via python demo.py against the real
Gemini API, including the cross-region/no-transport matching scenario.

Checklist

  • My code follows the project's style guidelines
  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published

Screenshots (if applicable)

N/A

Additional Notes

Built for the IDEALIZE 2026 mini-competition (AIESEC in University of Moratuwa / Yaala Labs).
Addresses SDG 11 (Sustainable Cities and Communities) and SDG 13 (Climate Action).

@Jenitha23
Jenitha23 requested a review from amithad as a code owner August 17, 2026 03:39

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove this file

@amithad amithad left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed against ak-dev-architecture, ak-dev-code-quality and ak-dev-testing-conventions. This is a self-contained use-case project under use-cases/, so the framework house patterns (pluggable factories, AKConfig knobs, classes-not-scripts) mostly don't bite here: the plain tool functions bound by OpenAIToolBuilder are the documented exception, and the process-global _STATE is explicitly justified in SPEC.md/README.md. Most findings are correctness and "does the documented path actually run" issues.

What's good

  • The agent/tool split is the right one: every judgement that must be consistent (urgency, match ranking, dedup) is deterministic Python, and the LLM only decides when to call it. SPEC.md states this and the code honours it.
  • 32 real unit tests on the deterministic layer, with an autouse state-reset fixture. Genuinely useful coverage, not smoke tests.
  • README.md/SPEC.md are honest about what is dummy data and what a production path looks like; the Limitations sections match the code.
  • WhatsApp dispatch is off by default and the tool result always says whether a send was real or simulated.

Findings: 4 blockers, 5 suggestions, 4 nits (inline). The three that matter most are all "the documented path does not run as written": the test-harness mode is invalid, .env.example is missing, and uv run pytest fails at collection.

Not anchorable to a diff line

  • Docs surfaces that enumerate use-cases/ were not updated. docs/docs/examples/overview.md (~L229) and docs/docs/agent-skills.md (~L176) both list the contents of use-cases/ and still name only waste-sorting-assistant. Adding a second use-case should add a bullet to each. (docs/src/pages/use-cases.tsx is a marketing page, not an inventory, so no change needed there.)
  • No CI ran on this PR. gh pr checks 628 reports no checks on the head branch. Note also that use-cases/ is outside EXAMPLE_DIRS in the Makefile, so make lint-check-all never formats this code. black/isort at line-length 120 would reshape a few spots (e.g. the implicit string concatenation at tool.py:61-62 and tool.py:590-591). Worth running uvx black -l 120 . && uvx isort --profile black -l 120 . in the project dir by hand.
  • Concurrency on _STATE. api.py runs through the queue pipeline with multiple agent-runner threads, and finalize_record's existing["quantity"] += ... is a read-modify-write on shared module state. Fine for a single-user demo; worth a line in Limitations alongside the existing "resets on restart" note if you want the caveat to be complete.

Already raised: @amithad's comment asking to remove diagnose_gemini.py still stands; not repeated inline.

#
# This project's agent replies are short and fairly predictable (see tests/test_agent_e2e.py's
# expect() lists), so fuzzy is enough and avoids needing a second LLM provider just for tests.
mode: fuzzy

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocker] fuzzy is not a valid test-harness mode, so this file fails validation and takes the whole e2e suite down with it.

  • AKTestConfig.mode is Field(default="fallback", pattern="^(fallback|llm|score)$") (ak-py/src/agentkernel/test/config.py:32). The three modes are score / llm / fallback, not fuzzy / judge / fallback.
  • Test.__init__ reads AKTestConfig.get().mode (ak-py/src/agentkernel/test/test.py:44), so constructing Test("demo.py") raises a pydantic ValidationError before a single request is sent.
  • This was never caught because all three test_agent_e2e.py tests skip without GEMINI_API_KEY: the PR description's "32 passed, 3 skipped" is exactly the run that cannot reach it. With a key set, the fixture errors.
  • Suggestion: mode: score, which is the deterministic, offline, no-extra-LLM-call mode the comment block above is actually describing. Then fix the comment (fuzzy→score, judge→llm) and the two places that repeat the wrong name: README.md:292 and SPEC.md:102.
  • Note score is exact-match-ish (quasi_exact_match_score), so the current expect([...]) keyword lists may need fallback instead once you can run it for real.

```bash
./build.sh
cp .env.example .env
# edit .env and set GEMINI_API_KEY (from https://aistudio.google.com/apikey)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocker] .env.example does not exist in this PR, so the documented first-run step fails.

  • git ls-tree on the PR head lists no .env.example; cp .env.example .env errors out, and it's step 2 of Setup.
  • Two other places promise the file exists: README.md:231 ("only .env.example (with no real values) is tracked in the repo") and the comment at agent.py:29.
  • Suggestion: add use-cases/disaster-response-agent/.env.example with exactly the keys listed at README.md:222-229 (GEMINI_API_KEY=, commented GEMINI_MODEL, WHATSAPP_ENABLED, AK_WHATSAPP__ACCESS_TOKEN, AK_WHATSAPP__PHONE_NUMBER_ID) and no real values. .gitignore already covers .env, so only the example gets committed.


import pytest

import tool

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocker] uv run pytest from the project root fails at collection, because tool is not importable from a tests/ subdirectory.

  • Reproduced locally with pytest 9.1.1 (the version this PR's uv.lock pins): ModuleNotFoundError: No module named 'tool' during collection of tests/test_probe.py.
  • Why: under pytest's default prepend import mode, the directory inserted into sys.path is the first ancestor of the test file without an __init__.py, which is tests/, not the project root. python -m pytest happens to work (it puts the cwd on sys.path), the pytest console script that uv run pytest invokes does not. So the commands in README.md:284/:290/:292 and in this file's own docstring all fail as written.
  • This also hits tests/test_agent_e2e.py, which imports agentkernel.test.Test("demo.py") against a path relative to the cwd.
  • Suggestion, in order of preference:
    1. Follow the repo convention: tests live beside demo.py as demo_test.py / tool_test.py (see examples/cli/openai/demo_test.py, examples/cli/custom-evaluator/demo_test.py, and the uv run pytest demo_test.py invocation in the bundled ak-test skill). No sys.path problem, and it matches every other AK project.
    2. Or keep tests/ and add to pyproject.toml:
      [tool.pytest.ini_options]
      pythonpath = ["."]

{
"id": "vol-001",
"name": "Nimal Perera",
"phone": "+94760048658",

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocker] A real personal phone number is committed to a public repository.

  • +94760048658 appears 9 times, on all six VOLUNTEER_DIRECTORY entries and all three seeded offers, and the comment at tool.py:92-98 confirms it is a real, WhatsApp-Cloud-API-verified number, not a placeholder.
  • yaalalabs/agent-kernel is public, so this lands in search indexes, clones and the published docs site. Anyone who runs the demo with WHATSAPP_ENABLED=true also sends live messages to it.
  • Suggestion: replace every occurrence with an obvious placeholder (+940000000000, or "" so _send_whatsapp_message returns its "No phone number on file" reason), and move the "add your own number as a verified test recipient in the Meta sandbox" instruction into README.md, where it's already most of the way there at README.md:165-167.

if c["resource_type"] == record["resource_type"] and c["status"] != "fulfilled"
]

requester_transport_flag = record.get("transport_flag") # True = need has NO transport

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Transport scoring assumes the intake is always a "need", so offer-initiated matching scores the worst pairing as the best.

  • transport_flag means two opposite things depending on message_type (_detect_transport_flag, L286-300): on a need True = "requester has no transport", on an offer True = "donor can deliver". This line reads it as the need meaning unconditionally, and match_resources runs in both directions (L556, offers→requests).
  • Verified against the PR head: submit the need "Need medicine in Matara, no transport", finalize it, then submit an offer "We have 30 medicine kits in Ratnapura" (donor cannot deliver). The match comes back match_score: 38, transport_note: "no transport constraint detected". The else branch at L592 awards the full 10 "no constraint" points to a stranded requester paired with a donor who cannot deliver, and tells dedup_dispatch_agent there is nothing to flag.
  • The elif ... and same_region branch (L585) is wrong the same way: for an offer intake it prints "requester has no transport" about the donor.
  • Suggestion: resolve the two roles from record["message_type"] before scoring, e.g.
    is_need = record["message_type"] == "need"
    requester_no_transport = record.get("transport_flag") if is_need else c.get("transport_flag")
    donor_can_deliver = c.get("transport_flag") if is_need else record.get("transport_flag")
    then keep the existing four-way branch on those two names. Worth a test mirroring test_no_transport_requester_with_no_compatible_donor_scores_lower but driven from the offer side.

readme = "README.md"
requires-python = ">=3.12"
dependencies = [
"agentkernel[cli,openai,api]>=0.6.1",

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] The agentkernel pin starts stale.

  • >=0.6.1 here vs >=0.8.1 in use-cases/waste-sorting-assistant/pyproject.toml, against a current release of 0.9.1.
  • scripts/update_examples_version.py only walks examples/ and e2e/app, so nothing bumps use-cases/ automatically, so whatever lands here is what stays until someone edits it by hand.
  • Suggestion: pin both this and the dev-group entry on L16 to the release you actually developed against, and re-run uv lock.


:param record_id: The id of the just-created/updated request or offer (from finalize_record).
:param matched_id: The id of the matched counterpart record to notify (from match_resources).
:param region: Unused - kept only for backward compatibility with older callers. Both

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] There are no "older callers": this is a brand-new file.

  • region is unused, but it's still in the tool schema OpenAIToolBuilder.bind generates, so the model is invited to fill a parameter that does nothing.
  • Suggestion: drop the parameter and the :param region: line; _find_record_by_id already searches every region.

:param region: The region/town to look up, e.g. "Galle".
:return: JSON string with open requests and offers for that region.
"""
store = _region_store(region)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] A read-only status query mutates the shared store.

  • _region_store ends in _STATE.setdefault(key, {...}) (L324), so asking "what's the status in Jaffna?" permanently adds an empty jaffna entry.
  • Harmless today, but it means _STATE accumulates a key per region anyone ever typed, and the region list stops meaning "regions with activity".
  • Suggestion: read without creating, e.g. store = _STATE.get(_normalize(region) or "unspecified", {"requests": {}, "offers": {}}).

@@ -0,0 +1,24 @@
"""Run the Disaster Response & Resource Coordination Agent locally via the Agent Kernel CLI.

Note: demo.py is the canonical entry point name used by Agent Kernel's other use-case

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] Two identical entry points is one to keep in sync for no gain.

  • cli.py and demo.py differ only in their docstrings; the code is the same four lines. Nothing in the repo has "muscle memory" for cli.py: use-cases/waste-sorting-assistant/ and every examples/** project ship demo.py alone.
  • Suggestion: delete cli.py and drop its mentions at README.md:301, README.md:372, SPEC.md:83 and the demo.py docstring.


- **SDG 11 - Sustainable Cities and Communities**: helps communities coordinate emergency
resources and respond to hazards faster, with less duplicated effort.
reduce disaster-related economic loss and disruption to essential services.

@amithad amithad Sep 21, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] Truncated sentence: a clause is missing before "reduce disaster-related economic loss".

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Multiple unresolved critical issues affect setup, test safety, credential handling, matching correctness, state consistency, and dispatch behavior.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 11 High severity · 2 Medium severity · 1 Low severity

Open (14)
What changed in this PR

Adds a Gemini/LiteLLM-powered three-agent disaster-response coordination system with matching, deduplication, dispatch, CLI/API entry points, tests, and documentation.

Changes:

  • Adds intake, priority/matching, deduplication, and WhatsApp dispatch tools.
  • Adds runtime configuration, build setup, diagnostics, and project documentation.
  • Adds deterministic unit tests and live conversational E2E tests.
File Review summary
use-cases/​disaster-response-agent/​tool.py Critical issues in transport matching, repeat dispatches, atomic deduplication, recipient fallback, and committed phone data. Moderate issues affect unit compatibility, failed notifications, and blocking I/O; a nit concerns stale tool documentation.
use-cases/​disaster-response-agent/​tests/​test_tool_layer.py Critical risk of sending real WhatsApp messages during tests; nit that the tests are not included in repository CI.
use-cases/​disaster-response-agent/​tests/​test_agent_e2e.py Moderate dotenv/skip-condition issue and critical fuzzy assertions that compare full responses with individual keywords.
use-cases/​disaster-response-agent/​test-config.yaml Critical fuzzy matching configuration does not correctly validate the E2E expectations.
use-cases/​disaster-response-agent/​SPEC.md Nit: persistence design should separate session storage from shared region-keyed disaster state.
use-cases/​disaster-response-agent/​README.md Critical setup instructions reference a missing .env.example file.
use-cases/​disaster-response-agent/​pyproject.toml Moderate dependency and lockfile alignment issue.
use-cases/​disaster-response-agent/​diagnose_gemini.py Critical credential exposure risk from unconditional debug logging; moderate model-default mismatch.
use-cases/​disaster-response-agent/​demo.py Reviewed; no final comments.
use-cases/​disaster-response-agent/​config.yaml Reviewed; no final comments.
use-cases/​disaster-response-agent/​cli.py Reviewed; no final comments.
use-cases/​disaster-response-agent/​build.sh Critical incorrect local wheel path and error masking.
use-cases/​disaster-response-agent/​api.py Reviewed; no final comments.
use-cases/​disaster-response-agent/​agent.py Reviewed; no final comments.
use-cases/​disaster-response-agent/​.python-version Reviewed; no final comments.
use-cases/​disaster-response-agent/​.gitignore Reviewed; no final comments.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.


```bash
./build.sh
cp .env.example .env
Comment on lines +14 to +15
uv sync --find-links ../agent-kernel/ak-py/dist --all-extras
uv pip install --force-reinstall --no-deps --no-index --find-links ../agent-kernel/ak-py/dist agentkernel[api,cli,openai,test] || true

import litellm

litellm._turn_on_debug() # print the raw HTTP request/response LiteLLM sends to Google
#
# This project's agent replies are short and fairly predictable (see tests/test_agent_e2e.py's
# expect() lists), so fuzzy is enough and avoids needing a second LLM provider just for tests.
mode: fuzzy
await test_client.send("Need drinking water in Galle")
# The seeded Galle offer should be matched same-region, so expect a positive confirmation
# mentioning water/Galle rather than a "nothing found" reply.
await test_client.expect(["water", "Galle", "recorded", "matched", "pending"])
Comment on lines +672 to +674
if duplicate_id and duplicate_id in store[pool_key]:
existing = store[pool_key][duplicate_id]
existing["quantity"] += record["quantity"]
Comment on lines +749 to +755
if not phone:
fallback = next(
(
v
for v in VOLUNTEER_DIRECTORY
if v["region"] == target_region and source["resource_type"] in v["resource_types"]
),
Comment on lines +17 to +23
import os

import pytest
import pytest_asyncio
from agentkernel.test import Test

pytestmark = [

scored = []
for c in candidates:
coverage = min(c["quantity"] / max(record["quantity"], 1), 1.0)
Comment on lines +69 to +73
- Production-ready path (not yet wired up, but the architecture is ready for it): swap `_STATE`
for Agent Kernel's Redis/DynamoDB/CosmosDB-backed session storage (see
`agent-kernel/examples/memory`), keyed by region instead of by session id, so the same shared
"disaster state" survives restarts and is visible to every process/instance handling traffic
for that disaster.
@amithad

amithad commented Sep 25, 2026

Copy link
Copy Markdown
Member

@Jenitha23 Please resolve the comments in the thread, if you interested in this use case being a part of Agent Kernel repository. Thank you

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants