Conversation
There was a problem hiding this comment.
Reviewed against ak-dev-architecture, ak-dev-code-quality and ak-dev-testing-conventions. This is a self-contained use-case project under use-cases/, so the framework house patterns (pluggable factories, AKConfig knobs, classes-not-scripts) mostly don't bite here: the plain tool functions bound by OpenAIToolBuilder are the documented exception, and the process-global _STATE is explicitly justified in SPEC.md/README.md. Most findings are correctness and "does the documented path actually run" issues.
What's good
- The agent/tool split is the right one: every judgement that must be consistent (urgency, match ranking, dedup) is deterministic Python, and the LLM only decides when to call it.
SPEC.mdstates this and the code honours it. - 32 real unit tests on the deterministic layer, with an autouse state-reset fixture. Genuinely useful coverage, not smoke tests.
README.md/SPEC.mdare honest about what is dummy data and what a production path looks like; the Limitations sections match the code.- WhatsApp dispatch is off by default and the tool result always says whether a send was real or simulated.
Findings: 4 blockers, 5 suggestions, 4 nits (inline). The three that matter most are all "the documented path does not run as written": the test-harness mode is invalid, .env.example is missing, and uv run pytest fails at collection.
Not anchorable to a diff line
- Docs surfaces that enumerate
use-cases/were not updated.docs/docs/examples/overview.md(~L229) anddocs/docs/agent-skills.md(~L176) both list the contents ofuse-cases/and still name onlywaste-sorting-assistant. Adding a second use-case should add a bullet to each. (docs/src/pages/use-cases.tsxis a marketing page, not an inventory, so no change needed there.) - No CI ran on this PR.
gh pr checks 628reports no checks on the head branch. Note also thatuse-cases/is outsideEXAMPLE_DIRSin theMakefile, somake lint-check-allnever formats this code.black/isortat line-length 120 would reshape a few spots (e.g. the implicit string concatenation attool.py:61-62andtool.py:590-591). Worth runninguvx black -l 120 . && uvx isort --profile black -l 120 .in the project dir by hand. - Concurrency on
_STATE.api.pyruns through the queue pipeline with multiple agent-runner threads, andfinalize_record'sexisting["quantity"] += ...is a read-modify-write on shared module state. Fine for a single-user demo; worth a line in Limitations alongside the existing "resets on restart" note if you want the caveat to be complete.
Already raised: @amithad's comment asking to remove diagnose_gemini.py still stands; not repeated inline.
| # | ||
| # This project's agent replies are short and fairly predictable (see tests/test_agent_e2e.py's | ||
| # expect() lists), so fuzzy is enough and avoids needing a second LLM provider just for tests. | ||
| mode: fuzzy |
There was a problem hiding this comment.
[blocker] fuzzy is not a valid test-harness mode, so this file fails validation and takes the whole e2e suite down with it.
AKTestConfig.modeisField(default="fallback", pattern="^(fallback|llm|score)$")(ak-py/src/agentkernel/test/config.py:32). The three modes arescore/llm/fallback, notfuzzy/judge/fallback.Test.__init__readsAKTestConfig.get().mode(ak-py/src/agentkernel/test/test.py:44), so constructingTest("demo.py")raises a pydanticValidationErrorbefore a single request is sent.- This was never caught because all three
test_agent_e2e.pytests skip withoutGEMINI_API_KEY: the PR description's "32 passed, 3 skipped" is exactly the run that cannot reach it. With a key set, the fixture errors. - Suggestion:
mode: score, which is the deterministic, offline, no-extra-LLM-call mode the comment block above is actually describing. Then fix the comment (fuzzy→score,judge→llm) and the two places that repeat the wrong name:README.md:292andSPEC.md:102. - Note
scoreis exact-match-ish (quasi_exact_match_score), so the currentexpect([...])keyword lists may needfallbackinstead once you can run it for real.
| ```bash | ||
| ./build.sh | ||
| cp .env.example .env | ||
| # edit .env and set GEMINI_API_KEY (from https://aistudio.google.com/apikey) |
There was a problem hiding this comment.
[blocker] .env.example does not exist in this PR, so the documented first-run step fails.
git ls-treeon the PR head lists no.env.example;cp .env.example .enverrors out, and it's step 2 of Setup.- Two other places promise the file exists:
README.md:231("only.env.example(with no real values) is tracked in the repo") and the comment atagent.py:29. - Suggestion: add
use-cases/disaster-response-agent/.env.examplewith exactly the keys listed atREADME.md:222-229(GEMINI_API_KEY=, commentedGEMINI_MODEL,WHATSAPP_ENABLED,AK_WHATSAPP__ACCESS_TOKEN,AK_WHATSAPP__PHONE_NUMBER_ID) and no real values..gitignorealready covers.env, so only the example gets committed.
|
|
||
| import pytest | ||
|
|
||
| import tool |
There was a problem hiding this comment.
[blocker] uv run pytest from the project root fails at collection, because tool is not importable from a tests/ subdirectory.
- Reproduced locally with pytest 9.1.1 (the version this PR's
uv.lockpins):ModuleNotFoundError: No module named 'tool'during collection oftests/test_probe.py. - Why: under pytest's default
prependimport mode, the directory inserted intosys.pathis the first ancestor of the test file without an__init__.py, which istests/, not the project root.python -m pytesthappens to work (it puts the cwd onsys.path), thepytestconsole script thatuv run pytestinvokes does not. So the commands inREADME.md:284/:290/:292and in this file's own docstring all fail as written. - This also hits
tests/test_agent_e2e.py, which importsagentkernel.test.Test("demo.py")against a path relative to the cwd. - Suggestion, in order of preference:
- Follow the repo convention: tests live beside
demo.pyasdemo_test.py/tool_test.py(seeexamples/cli/openai/demo_test.py,examples/cli/custom-evaluator/demo_test.py, and theuv run pytest demo_test.pyinvocation in the bundledak-testskill). Nosys.pathproblem, and it matches every other AK project. - Or keep
tests/and add topyproject.toml:[tool.pytest.ini_options] pythonpath = ["."]
- Follow the repo convention: tests live beside
| { | ||
| "id": "vol-001", | ||
| "name": "Nimal Perera", | ||
| "phone": "+94760048658", |
There was a problem hiding this comment.
[blocker] A real personal phone number is committed to a public repository.
+94760048658appears 9 times, on all sixVOLUNTEER_DIRECTORYentries and all three seeded offers, and the comment attool.py:92-98confirms it is a real, WhatsApp-Cloud-API-verified number, not a placeholder.yaalalabs/agent-kernelis public, so this lands in search indexes, clones and the published docs site. Anyone who runs the demo withWHATSAPP_ENABLED=truealso sends live messages to it.- Suggestion: replace every occurrence with an obvious placeholder (
+940000000000, or""so_send_whatsapp_messagereturns its "No phone number on file" reason), and move the "add your own number as a verified test recipient in the Meta sandbox" instruction intoREADME.md, where it's already most of the way there atREADME.md:165-167.
| if c["resource_type"] == record["resource_type"] and c["status"] != "fulfilled" | ||
| ] | ||
|
|
||
| requester_transport_flag = record.get("transport_flag") # True = need has NO transport |
There was a problem hiding this comment.
[suggestion] Transport scoring assumes the intake is always a "need", so offer-initiated matching scores the worst pairing as the best.
transport_flagmeans two opposite things depending onmessage_type(_detect_transport_flag, L286-300): on a needTrue= "requester has no transport", on an offerTrue= "donor can deliver". This line reads it as the need meaning unconditionally, andmatch_resourcesruns in both directions (L556, offers→requests).- Verified against the PR head: submit the need
"Need medicine in Matara, no transport", finalize it, then submit an offer"We have 30 medicine kits in Ratnapura"(donor cannot deliver). The match comes backmatch_score: 38,transport_note: "no transport constraint detected". Theelsebranch at L592 awards the full 10 "no constraint" points to a stranded requester paired with a donor who cannot deliver, and tellsdedup_dispatch_agentthere is nothing to flag. - The
elif ... and same_regionbranch (L585) is wrong the same way: for an offer intake it prints "requester has no transport" about the donor. - Suggestion: resolve the two roles from
record["message_type"]before scoring, e.g.then keep the existing four-way branch on those two names. Worth a test mirroringis_need = record["message_type"] == "need" requester_no_transport = record.get("transport_flag") if is_need else c.get("transport_flag") donor_can_deliver = c.get("transport_flag") if is_need else record.get("transport_flag")
test_no_transport_requester_with_no_compatible_donor_scores_lowerbut driven from the offer side.
| readme = "README.md" | ||
| requires-python = ">=3.12" | ||
| dependencies = [ | ||
| "agentkernel[cli,openai,api]>=0.6.1", |
There was a problem hiding this comment.
[suggestion] The agentkernel pin starts stale.
>=0.6.1here vs>=0.8.1inuse-cases/waste-sorting-assistant/pyproject.toml, against a current release of 0.9.1.scripts/update_examples_version.pyonly walksexamples/ande2e/app, so nothing bumpsuse-cases/automatically, so whatever lands here is what stays until someone edits it by hand.- Suggestion: pin both this and the dev-group entry on L16 to the release you actually developed against, and re-run
uv lock.
|
|
||
| :param record_id: The id of the just-created/updated request or offer (from finalize_record). | ||
| :param matched_id: The id of the matched counterpart record to notify (from match_resources). | ||
| :param region: Unused - kept only for backward compatibility with older callers. Both |
There was a problem hiding this comment.
[nit] There are no "older callers": this is a brand-new file.
regionis unused, but it's still in the tool schemaOpenAIToolBuilder.bindgenerates, so the model is invited to fill a parameter that does nothing.- Suggestion: drop the parameter and the
:param region:line;_find_record_by_idalready searches every region.
| :param region: The region/town to look up, e.g. "Galle". | ||
| :return: JSON string with open requests and offers for that region. | ||
| """ | ||
| store = _region_store(region) |
There was a problem hiding this comment.
[nit] A read-only status query mutates the shared store.
_region_storeends in_STATE.setdefault(key, {...})(L324), so asking "what's the status in Jaffna?" permanently adds an emptyjaffnaentry.- Harmless today, but it means
_STATEaccumulates a key per region anyone ever typed, and the region list stops meaning "regions with activity". - Suggestion: read without creating, e.g.
store = _STATE.get(_normalize(region) or "unspecified", {"requests": {}, "offers": {}}).
| @@ -0,0 +1,24 @@ | |||
| """Run the Disaster Response & Resource Coordination Agent locally via the Agent Kernel CLI. | |||
|
|
|||
| Note: demo.py is the canonical entry point name used by Agent Kernel's other use-case | |||
There was a problem hiding this comment.
[nit] Two identical entry points is one to keep in sync for no gain.
cli.pyanddemo.pydiffer only in their docstrings; the code is the same four lines. Nothing in the repo has "muscle memory" forcli.py:use-cases/waste-sorting-assistant/and everyexamples/**project shipdemo.pyalone.- Suggestion: delete
cli.pyand drop its mentions atREADME.md:301,README.md:372,SPEC.md:83and thedemo.pydocstring.
|
|
||
| - **SDG 11 - Sustainable Cities and Communities**: helps communities coordinate emergency | ||
| resources and respond to hazards faster, with less duplicated effort. | ||
| reduce disaster-related economic loss and disruption to essential services. |
There was a problem hiding this comment.
[nit] Truncated sentence: a clause is missing before "reduce disaster-related economic loss".
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Multiple unresolved critical issues affect setup, test safety, credential handling, matching correctness, state consistency, and dispatch behavior.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 11
Open (14)
Missing .env.example breaks fresh checkout setup · New Incorrect local wheel path breaks build · New Opt-in and redact LiteLLM debug logging · New Fuzzy tests incorrectly compare keywords as full responses · New Fuzzy assertions incorrectly compare keywords as full responses · New Prevent deterministic tests from sending real WhatsApp messages · New Remove committed real WhatsApp recipient number · New Exclude dispatched records from future matching · New Derive transport flags from intake and candidate message types · New Atomically deduplicate and finalize shared records · New Do not notify unrelated contacts for offer-to-need matches · New Load dotenv before evaluating live-test skip condition · New Validate units before computing quantity coverage · New Separate shared disaster state from session storage · New
What changed in this PR
Adds a Gemini/LiteLLM-powered three-agent disaster-response coordination system with matching, deduplication, dispatch, CLI/API entry points, tests, and documentation.
Changes:
- Adds intake, priority/matching, deduplication, and WhatsApp dispatch tools.
- Adds runtime configuration, build setup, diagnostics, and project documentation.
- Adds deterministic unit tests and live conversational E2E tests.
| File | Review summary |
|---|---|
use-cases/disaster-response-agent/tool.py |
Critical issues in transport matching, repeat dispatches, atomic deduplication, recipient fallback, and committed phone data. Moderate issues affect unit compatibility, failed notifications, and blocking I/O; a nit concerns stale tool documentation. |
use-cases/disaster-response-agent/tests/test_tool_layer.py |
Critical risk of sending real WhatsApp messages during tests; nit that the tests are not included in repository CI. |
use-cases/disaster-response-agent/tests/test_agent_e2e.py |
Moderate dotenv/skip-condition issue and critical fuzzy assertions that compare full responses with individual keywords. |
use-cases/disaster-response-agent/test-config.yaml |
Critical fuzzy matching configuration does not correctly validate the E2E expectations. |
use-cases/disaster-response-agent/SPEC.md |
Nit: persistence design should separate session storage from shared region-keyed disaster state. |
use-cases/disaster-response-agent/README.md |
Critical setup instructions reference a missing .env.example file. |
use-cases/disaster-response-agent/pyproject.toml |
Moderate dependency and lockfile alignment issue. |
use-cases/disaster-response-agent/diagnose_gemini.py |
Critical credential exposure risk from unconditional debug logging; moderate model-default mismatch. |
use-cases/disaster-response-agent/demo.py |
Reviewed; no final comments. |
use-cases/disaster-response-agent/config.yaml |
Reviewed; no final comments. |
use-cases/disaster-response-agent/cli.py |
Reviewed; no final comments. |
use-cases/disaster-response-agent/build.sh |
Critical incorrect local wheel path and error masking. |
use-cases/disaster-response-agent/api.py |
Reviewed; no final comments. |
use-cases/disaster-response-agent/agent.py |
Reviewed; no final comments. |
use-cases/disaster-response-agent/.python-version |
Reviewed; no final comments. |
use-cases/disaster-response-agent/.gitignore |
Reviewed; no final comments. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
|
||
| ```bash | ||
| ./build.sh | ||
| cp .env.example .env |
| uv sync --find-links ../agent-kernel/ak-py/dist --all-extras | ||
| uv pip install --force-reinstall --no-deps --no-index --find-links ../agent-kernel/ak-py/dist agentkernel[api,cli,openai,test] || true |
|
|
||
| import litellm | ||
|
|
||
| litellm._turn_on_debug() # print the raw HTTP request/response LiteLLM sends to Google |
| # | ||
| # This project's agent replies are short and fairly predictable (see tests/test_agent_e2e.py's | ||
| # expect() lists), so fuzzy is enough and avoids needing a second LLM provider just for tests. | ||
| mode: fuzzy |
| await test_client.send("Need drinking water in Galle") | ||
| # The seeded Galle offer should be matched same-region, so expect a positive confirmation | ||
| # mentioning water/Galle rather than a "nothing found" reply. | ||
| await test_client.expect(["water", "Galle", "recorded", "matched", "pending"]) |
| if duplicate_id and duplicate_id in store[pool_key]: | ||
| existing = store[pool_key][duplicate_id] | ||
| existing["quantity"] += record["quantity"] |
| if not phone: | ||
| fallback = next( | ||
| ( | ||
| v | ||
| for v in VOLUNTEER_DIRECTORY | ||
| if v["region"] == target_region and source["resource_type"] in v["resource_types"] | ||
| ), |
| import os | ||
|
|
||
| import pytest | ||
| import pytest_asyncio | ||
| from agentkernel.test import Test | ||
|
|
||
| pytestmark = [ |
|
|
||
| scored = [] | ||
| for c in candidates: | ||
| coverage = min(c["quantity"] / max(record["quantity"], 1), 1.0) |
| - Production-ready path (not yet wired up, but the architecture is ready for it): swap `_STATE` | ||
| for Agent Kernel's Redis/DynamoDB/CosmosDB-backed session storage (see | ||
| `agent-kernel/examples/memory`), keyed by region instead of by session id, so the same shared | ||
| "disaster state" survives restarts and is visible to every process/instance handling traffic | ||
| for that disaster. |
|
@Jenitha23 Please resolve the comments in the thread, if you interested in this use case being a part of Agent Kernel repository. Thank you |



Description
Adds a Gemini-powered submission for the IDEALIZE 2026 Agent Kernel mini-competition: a
three-agent Disaster Response & Resource Coordination system (intake → priority/matching →
dedup/dispatch) that parses free-form need/offer messages, scores urgency, matches across
regions with distance/transport awareness, deduplicates, and dispatches a WhatsApp notification.
Type of Change
Related Issues
Fixes #
Relates to #
Changes Made
use-cases/disaster-response-agent/: a 3-agent Agent Kernel pipeline(
intake_agent→priority_matching_agent→dedup_dispatch_agent) built on the OpenAIAgents SDK integration, using Gemini (
gemini-3.1-flash-lite) via LiteLLM's nativeintegration for all LLM calls
and transport compatibility ("no transport" / "can deliver" signals detected from message text)
WhatsApp dispatch notification, with automatic retry/backoff on rate-limited LLM calls
demo.py(canonical CLI entry point) andapi.py(REST API entry point)tests/test_tool_layer.py(32 deterministic unit tests) andtests/test_agent_e2e.py(live conversational test via
agentkernel.test.Test, skipped automatically without aGEMINI_API_KEY), plustest-config.yamlSPEC.mdandREADME.mdcovering the problem, architecture, agent/toolresponsibilities, memory design, SDG alignment, setup (including Windows-specific notes),
and limitations
Testing
Ran
uv run pytestlocally: 32 passed, 3 skipped (the live end-to-end tests, which require aGEMINI_API_KEY). Manually tested the full pipeline viapython demo.pyagainst the realGemini API, including the cross-region/no-transport matching scenario.
Checklist
Screenshots (if applicable)
N/A
Additional Notes
Built for the IDEALIZE 2026 mini-competition (AIESEC in University of Moratuwa / Yaala Labs).
Addresses SDG 11 (Sustainable Cities and Communities) and SDG 13 (Climate Action).