Conversation
* Updated config files and deployment scripts (#33) * Deployed prod MCP * Added ACM SSL cert * Updated shell script and staging vars * Staging tf vars files changed --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Generalized boston specific values --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix * Lint fix * bug fix --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix * Security update: removed Lambda entrypoint * Added DLQ SQS + tagging * Updated CLI tools for new service additions * Added transparency CLI commands + tests --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* lint fix * updated docs * Updated workflows * bug fix * Parallelized container setup * lint fix --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Removed redundant files * Github action to prevent divergent branching on PR to main * Added Socrata support (#27) * WIP: Socrata support (formatting etc) * Updated README file with right config * Updated Socrata isntructions to be more LLM friendly * Fixed smoke test bug * Updated smoke test bug --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Added SoSQL query support (#28) * Added SoSQL query support * Smoke test bug fix * Removed smoke test --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Updated docs * Feature/arcgis support (#30) * Bug fix * Ruff fix --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Socrata discovery API bug fix (#31) Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Pr/boston changes merge (#34) * Updated config files and deployment scripts (#33) * Deployed prod MCP * Added ACM SSL cert * Updated shell script and staging vars * Staging tf vars files changed --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Generalized boston specific values --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Feature/opencontext cli (#35) * Lint fix * Lint fix * bug fix --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Updated docs * Fixed CLI bugs and added template tfvars * lint fix * Added DX files * lint fix * Feature/security update (#37) * Lint fix * Security update: removed Lambda entrypoint * Added DLQ SQS + tagging * Updated CLI tools for new service additions * Added transparency CLI commands + tests --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Refactor/docs (thealphacubicle#61) * lint fix * updated docs * Updated workflows * bug fix * Parallelized container setup * lint fix --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Fixed CLI bug --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Updated config files and deployment scripts (#33) * Deployed prod MCP * Added ACM SSL cert * Updated shell script and staging vars * Staging tf vars files changed --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> * Generalized boston specific values --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix * Security update: removed Lambda entrypoint * Added DLQ SQS + tagging * Updated CLI tools for new service additions * Added transparency CLI commands + tests --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Fix issue template yamllint warnings * Claude support files * bug fix with terraform validation in CLI * Updated docs to be UV native * Docs fix + test case bug fix * Update .claude/rules/code-style.md Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * PR fixes * PR fixes --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
…cle#66) * Project health upgrades (thealphacubicle#64) * Fix issue template yamllint warnings * Claude support files * bug fix with terraform validation in CLI * Updated docs to be UV native * Docs fix + test case bug fix * Update .claude/rules/code-style.md Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * PR fixes * PR fixes --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Bulletin workflow added (thealphacubicle#65) Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local> --------- Co-authored-by: Srihari Raman <raman.sr@northeastern.edu> Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Added integration tests * chore: drop redundant files * chore: refactored AI-native context files * chore: fixed Claude specific AI docs bug * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* add initial GCP terraform and related artifacts * force specific version of typing-extensions * forcing typing-extensions downgrade * pin the correct max version of typing-extensions for Sentinel class support * align project.toml to requirements.txt * cli modifications for gcp authenticate/config/deploy/destroy/validate * add note re: gcp project id to readme * add initial GCP terraform and related artifacts * force specific version of typing-extensions * forcing typing-extensions downgrade * pin the correct max version of typing-extensions for Sentinel class support * align project.toml to requirements.txt * cli modifications for gcp authenticate/config/deploy/destroy/validate * add note re: gcp project id to readme * updated documentation * feat: gcp-related tests
server/lambda_handler.py duplicated server/adapters/aws_lambda.py with divergent behavior; it was dead code since the deployed handler is server.adapters.aws_lambda.lambda_handler (wired in terraform/aws/main.tf). The root local_server.py duplicated scripts/local_server.py, which is the maintained version (better logging, /mcp routing, OPENCONTEXT_CONFIG support). Docs updated to point at the maintained scripts/local_server.py. Co-Authored-By: Kimi K2.6 via opencode <noreply@ollama.com>
scripts/local_server.py serves MCP requests on /mcp (not /), so the FAQ and QUICKSTART curl examples targeting the root path would 404. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a provider plugin for Opendatasoft-based open data portals, built on
the shared decoupled base layer: OpendatasoftPlugin subclasses
BaseOpenDataPlugin (HTTP client lifecycle via _create_http_client, HTTP_RETRY,
_raise_http_error, ToolHandler dispatch with required-arg enforcement,
format_records/build_where_clause), and OpendatasoftPluginConfig subclasses
BasePluginConfig, reusing BasePluginConfig.validate_url as a staticmethod via
field_validator("base_url", "portal_url").
Toolset (6 tools): search_datasets, get_dataset, get_schema, query_data,
aggregate_data, list_categories — backed by the Explore v2.1 endpoints
/catalog/datasets, /catalog/datasets/{id}, /catalog/datasets/{id}/records and
/catalog/facets?facet=theme.
ODSQL validation: ODSQLValidator subclasses BaseQueryValidator and validates
where/select/order_by fragments by stripping single- and double-quoted string
literals before the shared forbidden-keyword scan, so keywords that appear
inside legitimate data values are not rejected. aggregate_data additionally
whitelists group_by fields and metric aliases with a safe-identifier regex and
metric expressions with a safe-aggregate regex covering count(*),
count(field), count(distinct field) and sum/avg/min/max(field), and accepts
the "field" | "-field" | "field ASC|DESC" order_by grammar (metric aliases
allowed, since ODSQL permits ordering by select aliases).
Long Beach (https://data.longbeach.gov) is the reference portal; all tests
mock HTTP and make no live network calls.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
(cherry picked from commit 6d1fbcd)
Adds an "Opendatasoft Plugin" section to docs/BUILT_IN_PLUGINS.md covering the configuration block, the six tools, ODSQL notes and validation behavior, the Explore v2.1 endpoints used, and Long Beach examples. Adds a commented, disabled opendatasoft block to config-example.yaml alongside the other built-in plugin examples. Co-Authored-By: Claude Opus <noreply@anthropic.com> (cherry picked from commit 5da76cc)
Live testing against data.longbeach.gov showed the Explore API repeats a group_by-less aggregate once per underlying record (100 identical rows for count(*)). Request limit=1 when no group_by is given — the single row is the whole answer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit b1e7888)
Seven confirmed findings fixed (the eighth, validator duplication across
plugins, is deferred to a core-level refactor):
- dataset_id was interpolated raw into the request path; a crafted id
('realid/exports/json?', '../facets') could redirect requests to other
endpoints. Ids are now validated as URL-slug identifiers at all three
call sites.
- The DataPlugin query_data path escaped string filters by doubling
single quotes (SQL convention) — an ODSQL syntax error for values like
"Val-d'Or". New _build_odsql_where emits double-quoted literals with
backslash escapes, ODSQL's convention.
- strip_literals blanked single-quoted spans before double-quoted ones,
so an apostrophe inside a double-quoted literal could hide structural
keywords from the forbidden-keyword scan. Both literal kinds are now
matched in one left-to-right pass.
- Limits are clamped to the API's 1..100 range on every path (query,
search, aggregate); previously limit=0/-1/500 produced empty results
or HTTP 400s.
- aggregate_data now coerces a bare-string group_by into a one-element
list instead of iterating it character by character.
- Query/aggregate formatters display every fetched record instead of
hardcoding a 10-row cap that discarded up to 90% of the transfer.
- count() with no argument now fails validation with a clear message
instead of reaching the API as an ODSQL syntax error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 50fabac)
All existing security work guards the outbound direction (LLM -> portal: SQL/SoQL validators, identifier whitelists, ArcGIS SSRF allow-list). Nothing guarded the inbound direction: every string a portal returned -- dataset descriptions, schema labels, error bodies, and record values -- was f-string-concatenated verbatim into tool results. Public datasets (311, permits, comments) contain text submitted by the public, so an attacker can plant instructions without compromising the portal, and hosts routinely pair this connector with tools that can act (email, files, calendar). New core/portal_content.py, applied centrally by BaseOpenDataPlugin: - Untrusted-data boundary: every successful text result is wrapped in a preamble + <<<BEGIN/END PORTAL DATA>>> markers; connector guidance moves to ToolHandler(guidance=...) and is emitted after the closing marker so instruction-shaped text never sits inside the data region. - Normalization: strip control/zero-width/bidi/tag/private-use code points, collapse newlines in single-line fields, explicit truncation markers, defang literal boundary markers; per-value, per-line, per-response caps. - Structure-forgery prevention: format_records keys are single-line and multi-line values/descriptions indent continuation lines so a value cannot fake a "Record N:" header or a connector hint at column 0. - safe_id + per-plugin id_pattern: IDs are only interpolated into Portal: links and hints if they match the provider's ID shape; links are built from config, never echoed from the portal. - ArcGIS _display_url: portal-supplied URLs shown only if the host passes the same allow-list that gates fetching. - Error bodies: _raise_http_error and JSON-RPC error.data cap/flatten portal text and label it "portal said:". - Heuristic injection detection: never blocks; prepends a WARNING line and logs a "Possible prompt injection markers" entry with the tool name for operator visibility. - MCP tool annotations: readOnlyHint/openWorldHint on every tool. Plugins (CKAN, Socrata, ArcGIS, Opendatasoft) route metadata through portal_line/portal_block/safe_id. docs/SECURITY.md documents the threat model and what the connector can and cannot guarantee. 37 new tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs (cherry picked from commit ee59322)
Socrata occasionally migrates a portal's domain (e.g. data.sfgov.org -> data.sf.gov) and 301s every path on the old one. httpx.AsyncClient defaults to not following redirects, so a portal_url that lags a rename got back a raw 301 with an HTML body: raise_for_status() didn't catch it (only 4xx/5xx), and response.json() failed trying to parse the redirect page. That broke get_schema/get_dataset/query_dataset/ execute_sql while search_datasets kept working, since the Discovery API treats old/new Socrata domains as aliases for search. Hit this live: data.sfgov.org started 301-redirecting to data.sf.gov, and every SODA3 tool failed until the deployment's portal_url was updated and this fix landed. Adds follow_redirects=True to the SODA client (the Discovery client is unaffected — api.us.socrata.com is stable infra, not a per-portal domain) plus a test asserting it's set. Claude-Session: https://claude.ai/code/session_01HuLvMhEAJgh6M1Rih5gWpD Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (cherry picked from commit 1f3b400)
…s everywhere (#20) Follow-up to #19. #19 made the Socrata SODA client follow redirects so a renamed portal domain (data.sfgov.org -> data.sf.gov) keeps working, but httpx only strips Authorization on cross-origin redirects — not custom headers such as Socrata's X-App-Token. A lapsed portal domain can be re-registered by someone else, who would then receive the token (and, for the other providers, any api key) from every deployment whose portal_url lags the rename. core/base_plugin.py: _create_http_client gains protect_headers= and trusted_hosts=. When set, it follows redirects and attaches a request event hook that drops those header names on any hop (initial or redirect) whose host is not the configured portal/base host, a subdomain of it, or a listed extra host. Shared _host_is_trusted / _trusted_request_hosts helpers. All four plugins opt in with their credential header: - Socrata SODA client: X-App-Token (replaces the bare follow_redirects=True) - CKAN, Opendatasoft: Authorization (also newly follow redirects) - ArcGIS hub + feature clients: Authorization, trusting *.arcgis.com and trusted_service_hosts (where feature services live) A legitimate rename is indistinguishable from a hijack, so the credential is not forwarded to the new host; the request still follows through unauthenticated (Socrata's app token is only a rate-limit key; public data on the others still resolves). Operators should update portal_url to the new domain. Tests: base hook coverage (same/subdomain kept, untrusted dropped + logged, rename drops, extra trusted host retained, follow_redirects forced) and a Socrata assertion that the SODA client opts in. docs/SECURITY.md documents the scoping. 624 tests pass; no new ruff findings vs main. Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit 260fbc0)
…#14) * feat(ckan): surface catalog metadata and add list_datasets + get_catalog_stats A Boston eval comparing an agent with and without OpenContext flagged the connector's top gap: get_dataset summarized package_show into prose and dropped metadata_modified, license, organization, tags, and every resource's url/created/last_modified (the resource URL alone would have dated the year-series resources the agent had to leave null), and there was no way to count the catalog, so catalog-wide statistics were asserted instead of counted. core/base_plugin.py (shared by all plugins): - short_date (ISO / epoch s / epoch ms -> YYYY-MM-DD), human_size, format_search_header (catalog-wide total + "showing a-b"), and display_portal_url: a portal-supplied URL is echoed only when its host is the portal/API host or a subdomain (or an extra trusted host); otherwise only "(external: hostname)" is shown. Generalizes ArcGIS _display_url. plugins/ckan/plugin.py: - get_dataset now prints organization (title + slug), license (title + id), created/modified, tags, groups, and per resource: created, modified (last_modified or metadata_modified), size, DataStore flag, gated download URL and description. New max_resources arg (default 50, max 500) with an explicit "... and N more" line. - search_datasets header uses package_search's count; rows show organization, modified date, resource count + formats, tag count. - New list_datasets: exact-match filters (organization/tag/format/license/ group), sort enum (default metadata_modified desc), limit/offset, total count. Filters are whitelisted and values quoted as escaped Solr phrases (_build_fq), so model input cannot alter the fq query. fl is deliberately not used (it collapses organization and drops resources). - New get_catalog_stats: package_search rows=0 + facet.field for organization/tags/res_format/license_id/groups, optional query/filters, values sorted by count; falls back to legacy `facets` on older CKAN. - Guidance strings point the model between the three catalog tools. Tests: 22 new (Solr escaping/whitelist, enriched formatters incl. external and non-http resource URLs, max_resources clamp/truncation, count header, list_datasets request shape and invalid sort, catalog stats request shape, sorting, legacy fallback, empty facets, unknown facet). Docs updated (BUILT_IN_PLUGINS, ARCHITECTURE incl. previously missing aggregate_data, SECURITY). Verified live against data.boston.gov: 235 datasets, facet buckets, boston-311-org listing sorted by modified, 311 dataset with 21 dated resources. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs * style: datetime.UTC, ClassVar test fixtures, drop unused noqa Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs * style: ClassVar on remaining test fixtures Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs * style: hoist CKAN helper imports to module top (CI ruff E402) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit 384b3d0)
…ta, ODS, ArcGIS (#16) Companion to the CKAN catalog work (#14): the same shape of loss existed in the other three providers — none passed the search total to its formatter (Discovery resultSetSize / ODS total_count / Hub numberMatched), and each get_dataset dropped dates, license, publisher/attribution and counts the API had already returned. Socrata: _discovery_search returns the envelope; search header uses resultSetSize and rows show modified date, column count, downloads, source. get_dataset adds source/attribution (link host-gated), license name + id (terms link host-gated) instead of a stringified dict, created/published/ metadata-modified/rows-updated dates, rows/columns/downloads/views/ provenance; empty values are omitted instead of "N/A". Opendatasoft: _catalog_search returns the envelope; header uses total_count (the pattern _tool_query_data already used) and rows show publisher, modified, records. get_dataset adds publisher, license (+ gated URL), attribution, data/metadata processed dates, field count, references. ArcGIS: _search_hub returns results + numberMatched; rows show owner, created/modified, record count. get_dataset adds organization, last-edit date, size, record count, categories, type keywords, access information; empty fields are omitted. _display_url now delegates to the shared display_portal_url with *.arcgis.com + trusted_service_hosts, so untrusted hosts render as "(external: host)" like every other provider. Tests: base helper coverage (short_date, human_size, display_portal_url, format_search_header) plus enrichment tests per provider. Docs updated. Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit 0ae1dcb)
Socrata's developer console issues two different credential types:
a single-string App Token (Developer Settings -> App Tokens), and an
API Key pair (Key ID + Key Secret, Developer Settings -> API Keys)
meant for HTTP Basic Auth on authenticated requests. This plugin only
implements the former, sending app_token bare as X-App-Token.
Pasting an API Key's Key ID in as app_token fails silently for some
tools and not others: search_datasets/get_dataset keep working
(catalog/metadata calls tolerate it), but query_dataset fails with
"Invalid app_token specified" (403) on /resource/{id}.json — the
actual SODA3 data-query endpoint. Confirmed against data.ny.gov,
data.cityofnewyork.us, data.lacity.org, and data.sfgov.org while
deploying new portals.
Updates the config schema field description/validation error and
BUILT_IN_PLUGINS.md to steer setup toward the right credential type.
No auth code changes needed — a bare App Token is sufficient since
Socrata open-data portals are all public.
Claude-Session: https://claude.ai/code/session_01HuLvMhEAJgh6M1Rih5gWpD
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit e1ff187)
… _search_hub tool path; ruff format
…CP adapter, test layout Adopts thealphacubicle/OpenContext develop as the fork's baseline, carrying every fork improvement already ported there (shared base layer, Opendatasoft, portal-content guardrails, redirect credential scoping, catalog metadata). Brings in the Typer CLI (opencontext configure/deploy/serve/...), the GCP Cloud Functions adapter and terraform, the unit/security/integration/smoke test layout with the upstream suites, and upstream docs and AI context. Drops the Boston-specific deploy script and tfvars; deployment is now driven by the CLI from a generated config. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
- sqlparse>=0.6.0, aiohttp>=3.14.3, pytest>=9.0.3 (Dependabot bumps the fork had already taken); uv.lock re-resolved. - CI test job runs on Python 3.11 and 3.12 with upstream's coverage gate. - Drop the sync-develop workflow: this fork has no develop branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
- Local dev servers (`opencontext serve`, tests' local_server) no longer echo exception text to the client; details stay in the log (py/stack-trace-exposure). - Tests compare URL sets / use anchored regexes instead of substring checks CodeQL reads as URL sanitization (py/incomplete-url-substring-sanitization). - Workflows declare `permissions: contents: read` at the top level (actions/missing-workflow-permissions); release.yml jobs keep their own write grants. - Drop notify-bulletin.yml: it is gated to thealphacubicle/OpenContext and posts to their bulletin endpoint, so it never runs here. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
Adopt thealphacubicle develop as baseline: CLI, GCP adapter, test layout
… f=json, schema fallback, optional app token (#26) * fix(arcgis): auto-trust Hub-referenced service hosts; schema fallback; wizard prompt Findings from the 51-jurisdiction open-data audit (Open_Data_Inventory, reports/gaps_and_improvements.md §4.5, §4.10) and thealphacubicle PR thealphacubicle#75: - Hub catalogs routinely reference Feature Services on city GIS domains (gis.charlottenc.gov, maps2.dcgis.dc.gov, gis.indy.gov). The SSRF guard refused them unless each host was pre-listed, so Charlotte, Denver and Indianapolis lost query_data/get_schema entirely. New `auto_trust_hub_services` (default true) accepts a Hub-referenced URL when it is https, on a public DNS name (IP literals, single labels and .internal/.local names are always refused), and has an ArcGIS REST service path. The bearer token is never sent to auto-trusted hosts (protect_headers already scopes it to portal/arcgis.com/allow-list). Explicit trusted_service_hosts still wins, and operators can turn auto-trust off. - Refusals are structured: the error starts with `untrusted_service_host: '<host>'` and names the config key, instead of surfacing as a generic "not trusted"/JSON-decode failure that three audits misdiagnosed. - get_schema derives the field list from a one-row /query when the layer metadata endpoint returns non-JSON or an ArcGIS error envelope (Columbus, Indianapolis: metadata endpoint flaky while /query worked). - Service URLs are whitespace-stripped before validation. - `opencontext configure` prompts for trusted_service_hosts in the ArcGIS step (ported from PR thealphacubicle#75 by cstirry, mapped onto this field name). Docs: BUILT_IN_PLUGINS.md, SECURITY.md, ARCHITECTURE.md, config-example.yaml. Tests: auto-trust acceptance/refusal matrix, plugin-level on/off, schema fallback paths, CLI prompt parsing. 1105 passed, 94% coverage, ruff clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX * fix(plugins): port #23 — f=json as a param, first-layer resolution, optional Socrata token Ports #23 (fix/connector-gaps-cdc-hud, found against data.cdc.gov and HUD's ArcGIS Hub, 2026-09-20/21) onto the new layout, merged with the auto-trust and schema-fallback work on this branch: - ArcGIS get_schema built `{url}/0?f=json` and passed `params={}`; httpx replaces the URL's query string with params, so the request went out bare, ArcGIS answered HTML, and `.json()` failed on every dataset. This is the root cause behind the "JSON-decode" schema failures the audit logged for Columbus and Denver. The format now travels as `params={"f": "json"}`; the one-row-query fallback stays for services whose metadata endpoint is genuinely broken, and reports a ValueError naming both endpoints when the query path is unusable too. - `_resolve_layer_url` reads the service description once, uses the first layer id (then the first table, then 0), caches per service, and keeps an explicit `/N`. Services whose only layer is not id 0 (HUD Low-Mod Income by Tract = layer 4, Opportunity Zones = layer 13) failed with "Invalid URL". get_schema and _query_features both use it. - Socrata `app_token` is optional: blank means no X-App-Token header. data.cdc.gov serves SODA3 untokened but rejects an invalid token (403), so a deployer without a token could not run the plugin at all (audit §4.1). Docs updated. Tests from #23 land under tests/unit/plugins/{arcgis,socrata}/. 1111 passed, 94% coverage, ruff clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX * ci: run the Terraform check on every PR so the required status always reports Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Bumps [anyio](https://github.com/agronholm/anyio) from 4.12.0 to 4.14.2. - [Release notes](https://github.com/agronholm/anyio/releases) - [Commits](agronholm/anyio@4.12.0...4.14.2) --- updated-dependencies: - dependency-name: anyio dependency-version: 4.14.2 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [starlette](https://github.com/Kludex/starlette) from 0.52.1 to 1.3.1. - [Release notes](https://github.com/Kludex/starlette/releases) - [Changelog](https://github.com/Kludex/starlette/blob/main/docs/release-notes.md) - [Commits](Kludex/starlette@0.52.1...1.3.1) --- updated-dependencies: - dependency-name: starlette dependency-version: 1.3.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
format_records defaulted to max_display=10, so query_data on ArcGIS (and the CKAN/Socrata SQL and aggregate paths, which hard-coded 10, and CKAN query_data, which hard-coded 5) fetched up to `limit` rows but showed only the first few. Clients re-paged for rows they had already received. max_display now defaults to None (all records). A character budget just under the framing cap drops whole records from the end and says "Showing N of M record(s)" instead of cutting a record mid-value. Closes #28 Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Resolves the open Dependabot alerts on uv.lock and one advisory pip-audit reports that Dependabot has not raised yet: - urllib3 2.6.3 -> 2.8.0 (alerts #16, #17, high) - msgpack 1.1.2 -> 1.2.2 (alert #30, high) - idna 3.11 -> 3.20 (alert #18) - pip 26.0.1 -> 26.2.1 (alerts #14, #15, #31, #39; dev-only, via pip-audit) - click 8.3.1 -> 8.3.3 (PYSEC-2026-2132; 8.5 deprecates APIs typer uses) Lockfile only; no pyproject or requirements.txt changes. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
query_data gains `offset` (resultOffset), `order_by` (orderByFields, validated to field names with optional ASC/DESC) and `format` (text | json | csv). Replies now start with "Returned N of M matching record(s) (offset, limit)" and end with "Next page: offset=..." when more records match. The total comes from a returnCountOnly request, skipped when a short first page already proves it; a failed count falls back to exceededTransferLimit. The shared base gains render_rows(), which renders text blocks, a JSON array (one object per line) or CSV within the response budget, dropping whole rows and reporting how many were shown so the next offset is exact. format_records is now a wrapper over it. ArcGIS plugin_version 1.0.0 -> 1.1.0. Closes #29 Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
) A Feature Service can hold several layers and tables, but get_schema and query_data only reached the default one. HUD's Small Area FMRs keep the rents in table 1 behind a geometry-only layer 0, so they were unreachable. - get_schema and query_data take an optional integer `layer`. The id is checked against the service's layers and tables; an unknown id is refused with the list of valid ones. - get_dataset lists the service's layers and tables (id, name, geometry or "table") and marks the default. The lookup is skipped for untrusted hosts and never fails the call. - The service description is read once per service and cached; the default-layer resolution now reuses it. Closes #30 Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Totals by area needed every row pulled and summed client-side: five-county
voucher sums took 23 paged query_data calls. aggregate_data sends one
outStatistics query with groupByFieldsForStatistics instead.
- statistics: [{type: count|sum|avg|min|max|stddev, field, as}]; a count
without a field counts records by the object-id field.
- group_by, where, having, order_by, limit, layer, format (text|json|csv).
- Field and group names must be plain identifiers and exist in the layer
schema (case-insensitive; sent with the schema's spelling). where and
having use WhereValidator; order_by uses validate_order_by. All input
checks run before any request.
- Layers reporting supportsStatistics: false are refused.
- The guidance line says nulls are excluded, so sums over suppressed
values are lower bounds.
get_schema, query_data and aggregate_data now share _layer_url_for
(item -> trusted layer URL) and _layer_fields (fields + metadata, with
the one-row-query fallback).
Closes #31
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…38) Descriptions were cut at 300 characters of raw HTML, hiding definitions and suppression rules; the only date shown was the Hub item's; answers carried nothing to cite. - core/portal_content.html_to_text (stdlib html.parser): block tags to lines, list bullets, entities decoded, script/style dropped. Applied to item descriptions, snippets, licence and access information. - get_dataset shows the whole description (capped at 12,000 characters with a truncation notice); search keeps a 300-character excerpt. - get_dataset adds the default layer's data and schema edit dates from editingInfo (HUD FMR: 2025-09-30). Layer descriptions are cached. - get_schema adds string field lengths and coded-value/range domains. - query_data and aggregate_data end with a source block: dataset title and Hub page, layer queried, data edit date (only from the cache, so no extra request), licence, retrieval time. ArcGIS metadata XML (field labels) is left for a follow-up. Closes #32 Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- ARCHITECTURE.md: the ArcGIS tool table listed four tools with their old arguments; it now lists get_schema and aggregate_data, the layer, paging and format arguments, and links to BUILT_IN_PLUGINS.md. - BUILT_IN_PLUGINS.md: search/aggregation tools take `query`, not `q`; replace the "appends /0" note with first-layer resolution and `layer`; schema fallback path uses the resolved layer; WHERE validation covers having and order_by; get_schema row mentions lengths and domains. - CUSTOM_PLUGINS.md: document render_rows. - SECURITY.md: ArcGIS order_by / aggregate_data / layer input checks, html_to_text, and render_rows in the structure-forgery row. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR brings over the plugin-architecture refactor plus follow-up review fixes and CI hardening (30 commits, fast-forward from
CityOfBoston:main).Shared plugin architecture
core/base_plugin.py(HTTP lifecycle, retry, tool dispatch, record formatting,build_where_clausewith identifier validation),core/config_base.py,core/query_validator.pyserver/lambda_handler.pyand rootlocal_server.pyreplaced byserver/adapters/aws_lambda.py(persistent event loop — no more per-invocationasyncio.run) andscripts/local_server.pySecurity hardening
*.arcgis.com, portal host, plus configurabletrusted_service_hosts)build_where_clause(injection through field names now fails closed for all plugins, including third-party)aggregate_datavalidates metrics, HAVING keys/values, andorder_byagainst safe whitelistsstatus = 'SET'still workReview-driven fixes (verified by an 8-finding code review)
order_bysupportsfield DESC/-fieldcorrectlycount(field)/count(distinct field)and metric-alias HAVING keys acceptedqueryargument enforced on search tools (ArcGIS, Socrata)validate_urlmade a staticmethod (pydantic 2.12 compatibility)CI and dependency hygiene
terraform/awsandterraform/bootstrap)pip-auditCVE scan and ruff already enforced; coverage gate at 80%Testing
terraform validateclean for both stacksNotes for reviewers
mainsearch_datasets(q→query, now required)🤖 Generated with Claude Code