Skip to content

Plugin architecture refactor, hardened validation, and CI enforcement - #1

Draft
sgarcese wants to merge 132 commits into
CityOfBoston:mainfrom
sgarcese:main
Draft

sgarcese wants to merge 132 commits into
CityOfBoston:mainfrom
sgarcese:main

Conversation

@sgarcese

Copy link
Copy Markdown

Summary

This PR brings over the plugin-architecture refactor plus follow-up review fixes and CI hardening (30 commits, fast-forward from CityOfBoston:main).

Shared plugin architecture

  • New shared base classes: core/base_plugin.py (HTTP lifecycle, retry, tool dispatch, record formatting, build_where_clause with identifier validation), core/config_base.py, core/query_validator.py
  • ArcGIS, CKAN, and Socrata plugins migrated onto the shared bases; plugin template modernized
  • server/lambda_handler.py and root local_server.py replaced by server/adapters/aws_lambda.py (persistent event loop — no more per-invocation asyncio.run) and scripts/local_server.py

Security hardening

  • SSRF guard for ArcGIS Feature Service URLs (allowlist: *.arcgis.com, portal host, plus configurable trusted_service_hosts)
  • Field-identifier validation inside build_where_clause (injection through field names now fails closed for all plugins, including third-party)
  • CKAN aggregate_data validates metrics, HAVING keys/values, and order_by against safe whitelists
  • Forbidden-keyword scan skips quoted string literals, so legitimate values like status = 'SET' still work

Review-driven fixes (verified by an 8-finding code review)

  • order_by supports field DESC / -field correctly
  • count(field) / count(distinct field) and metric-alias HAVING keys accepted
  • Required query argument enforced on search tools (ArcGIS, Socrata)
  • ArcGIS results use the shared capped record formatter
  • validate_url made a staticmethod (pydantic 2.12 compatibility)

CI and dependency hygiene

  • Test suite matrixed across Python 3.11 / 3.12
  • New Terraform validate job (fmt + validate for terraform/aws and terraform/bootstrap)
  • Dependabot config (pip, gomod, github-actions, terraform)
  • pip-audit CVE scan and ruff already enforced; coverage gate at 80%

Testing

Notes for reviewers

  • Recommend enabling branch protection with the four CI checks (Code quality, Test suite 3.11/3.12, Terraform validate) as required status checks on main
  • Breaking-ish behavior changes vs the pre-refactor code are listed in the security section above; tool schemas changed for ArcGIS search_datasets (q → query, now required)

🤖 Generated with Claude Code

thealphacubicle and others added 30 commits March 31, 2026 12:20
* Updated config files and deployment scripts (#33)

* Deployed prod MCP

* Added ACM SSL cert

* Updated shell script and staging vars

* Staging tf vars files changed

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Generalized boston specific values

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix

* Lint fix

* bug fix

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix

* Security update: removed Lambda entrypoint

* Added DLQ SQS + tagging

* Updated CLI tools for new service additions

* Added transparency CLI commands + tests

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* lint fix

* updated docs

* Updated workflows

* bug fix

* Parallelized container setup

* lint fix

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Removed redundant files

* Github action to prevent divergent branching on PR to main

* Added Socrata support  (#27)

* WIP: Socrata support (formatting etc)

* Updated README file with right config

* Updated Socrata isntructions to be more LLM friendly

* Fixed smoke test bug

* Updated smoke test bug

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Added SoSQL query support (#28)

* Added SoSQL query support

* Smoke test bug fix

* Removed smoke test

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Updated docs

* Feature/arcgis support (#30)

* Bug fix

* Ruff fix

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Socrata discovery API bug fix (#31)

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Pr/boston changes merge (#34)

* Updated config files and deployment scripts (#33)

* Deployed prod MCP

* Added ACM SSL cert

* Updated shell script and staging vars

* Staging tf vars files changed

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Generalized boston specific values

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Feature/opencontext cli (#35)

* Lint fix

* Lint fix

* bug fix

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Updated docs

* Fixed CLI bugs and added template tfvars

* lint fix

* Added DX files

* lint fix

* Feature/security update (#37)

* Lint fix

* Security update: removed Lambda entrypoint

* Added DLQ SQS + tagging

* Updated CLI tools for new service additions

* Added transparency CLI commands + tests

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Refactor/docs (thealphacubicle#61)

* lint fix

* updated docs

* Updated workflows

* bug fix

* Parallelized container setup

* lint fix

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Fixed CLI bug

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Updated config files and deployment scripts (#33)

* Deployed prod MCP

* Added ACM SSL cert

* Updated shell script and staging vars

* Staging tf vars files changed

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>

* Generalized boston specific values

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Lint fix

* Security update: removed Lambda entrypoint

* Added DLQ SQS + tagging

* Updated CLI tools for new service additions

* Added transparency CLI commands + tests

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
* Fix issue template yamllint warnings

* Claude support files

* bug fix with terraform validation in CLI

* Updated docs to be UV native

* Docs fix + test case bug fix

* Update .claude/rules/code-style.md

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* PR fixes

* PR fixes

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
…cle#66)

* Project health upgrades (thealphacubicle#64)

* Fix issue template yamllint warnings

* Claude support files

* bug fix with terraform validation in CLI

* Updated docs to be UV native

* Docs fix + test case bug fix

* Update .claude/rules/code-style.md

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* PR fixes

* PR fixes

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Bulletin workflow added (thealphacubicle#65)

Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>

---------

Co-authored-by: Srihari Raman <raman.sr@northeastern.edu>
Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Added integration tests

* chore: drop redundant files

* chore: refactored AI-native context files

* chore: fixed Claude specific AI docs bug

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Srihari Raman <srihariraman@Sriharis-MacBook-Pro.local>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* add initial GCP terraform and related artifacts

* force specific version of typing-extensions

* forcing typing-extensions downgrade

* pin the correct max version of typing-extensions for Sentinel class support

* align project.toml to requirements.txt

* cli modifications for gcp authenticate/config/deploy/destroy/validate

* add note re: gcp project id to readme

* add initial GCP terraform and related artifacts

* force specific version of typing-extensions

* forcing typing-extensions downgrade

* pin the correct max version of typing-extensions for Sentinel class support

* align project.toml to requirements.txt

* cli modifications for gcp authenticate/config/deploy/destroy/validate

* add note re: gcp project id to readme

* updated documentation

* feat: gcp-related tests
server/lambda_handler.py duplicated server/adapters/aws_lambda.py with divergent behavior; it was dead code since the deployed handler is server.adapters.aws_lambda.lambda_handler (wired in terraform/aws/main.tf).

The root local_server.py duplicated scripts/local_server.py, which is the maintained version (better logging, /mcp routing, OPENCONTEXT_CONFIG support).

Docs updated to point at the maintained scripts/local_server.py.

Co-Authored-By: Kimi K2.6 via opencode <noreply@ollama.com>
scripts/local_server.py serves MCP requests on /mcp (not /), so the
FAQ and QUICKSTART curl examples targeting the root path would 404.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
sgarcese and others added 30 commits September 13, 2026 17:03
Adds a provider plugin for Opendatasoft-based open data portals, built on
the shared decoupled base layer: OpendatasoftPlugin subclasses
BaseOpenDataPlugin (HTTP client lifecycle via _create_http_client, HTTP_RETRY,
_raise_http_error, ToolHandler dispatch with required-arg enforcement,
format_records/build_where_clause), and OpendatasoftPluginConfig subclasses
BasePluginConfig, reusing BasePluginConfig.validate_url as a staticmethod via
field_validator("base_url", "portal_url").

Toolset (6 tools): search_datasets, get_dataset, get_schema, query_data,
aggregate_data, list_categories — backed by the Explore v2.1 endpoints
/catalog/datasets, /catalog/datasets/{id}, /catalog/datasets/{id}/records and
/catalog/facets?facet=theme.

ODSQL validation: ODSQLValidator subclasses BaseQueryValidator and validates
where/select/order_by fragments by stripping single- and double-quoted string
literals before the shared forbidden-keyword scan, so keywords that appear
inside legitimate data values are not rejected. aggregate_data additionally
whitelists group_by fields and metric aliases with a safe-identifier regex and
metric expressions with a safe-aggregate regex covering count(*),
count(field), count(distinct field) and sum/avg/min/max(field), and accepts
the "field" | "-field" | "field ASC|DESC" order_by grammar (metric aliases
allowed, since ODSQL permits ordering by select aliases).

Long Beach (https://data.longbeach.gov) is the reference portal; all tests
mock HTTP and make no live network calls.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
(cherry picked from commit 6d1fbcd)
Adds an "Opendatasoft Plugin" section to docs/BUILT_IN_PLUGINS.md covering the
configuration block, the six tools, ODSQL notes and validation behavior, the
Explore v2.1 endpoints used, and Long Beach examples. Adds a commented,
disabled opendatasoft block to config-example.yaml alongside the other
built-in plugin examples.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
(cherry picked from commit 5da76cc)
Live testing against data.longbeach.gov showed the Explore API repeats
a group_by-less aggregate once per underlying record (100 identical
rows for count(*)). Request limit=1 when no group_by is given — the
single row is the whole answer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b1e7888)
Seven confirmed findings fixed (the eighth, validator duplication across
plugins, is deferred to a core-level refactor):

- dataset_id was interpolated raw into the request path; a crafted id
  ('realid/exports/json?', '../facets') could redirect requests to other
  endpoints. Ids are now validated as URL-slug identifiers at all three
  call sites.
- The DataPlugin query_data path escaped string filters by doubling
  single quotes (SQL convention) — an ODSQL syntax error for values like
  "Val-d'Or". New _build_odsql_where emits double-quoted literals with
  backslash escapes, ODSQL's convention.
- strip_literals blanked single-quoted spans before double-quoted ones,
  so an apostrophe inside a double-quoted literal could hide structural
  keywords from the forbidden-keyword scan. Both literal kinds are now
  matched in one left-to-right pass.
- Limits are clamped to the API's 1..100 range on every path (query,
  search, aggregate); previously limit=0/-1/500 produced empty results
  or HTTP 400s.
- aggregate_data now coerces a bare-string group_by into a one-element
  list instead of iterating it character by character.
- Query/aggregate formatters display every fetched record instead of
  hardcoding a 10-row cap that discarded up to 90% of the transfer.
- count() with no argument now fails validation with a clear message
  instead of reaching the API as an ODSQL syntax error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 50fabac)
All existing security work guards the outbound direction (LLM -> portal:
SQL/SoQL validators, identifier whitelists, ArcGIS SSRF allow-list). Nothing
guarded the inbound direction: every string a portal returned -- dataset
descriptions, schema labels, error bodies, and record values -- was
f-string-concatenated verbatim into tool results. Public datasets (311,
permits, comments) contain text submitted by the public, so an attacker can
plant instructions without compromising the portal, and hosts routinely pair
this connector with tools that can act (email, files, calendar).

New core/portal_content.py, applied centrally by BaseOpenDataPlugin:

- Untrusted-data boundary: every successful text result is wrapped in a
  preamble + <<<BEGIN/END PORTAL DATA>>> markers; connector guidance moves to
  ToolHandler(guidance=...) and is emitted after the closing marker so
  instruction-shaped text never sits inside the data region.
- Normalization: strip control/zero-width/bidi/tag/private-use code points,
  collapse newlines in single-line fields, explicit truncation markers,
  defang literal boundary markers; per-value, per-line, per-response caps.
- Structure-forgery prevention: format_records keys are single-line and
  multi-line values/descriptions indent continuation lines so a value cannot
  fake a "Record N:" header or a connector hint at column 0.
- safe_id + per-plugin id_pattern: IDs are only interpolated into Portal:
  links and hints if they match the provider's ID shape; links are built
  from config, never echoed from the portal.
- ArcGIS _display_url: portal-supplied URLs shown only if the host passes
  the same allow-list that gates fetching.
- Error bodies: _raise_http_error and JSON-RPC error.data cap/flatten portal
  text and label it "portal said:".
- Heuristic injection detection: never blocks; prepends a WARNING line and
  logs a "Possible prompt injection markers" entry with the tool name for
  operator visibility.
- MCP tool annotations: readOnlyHint/openWorldHint on every tool.

Plugins (CKAN, Socrata, ArcGIS, Opendatasoft) route metadata through
portal_line/portal_block/safe_id. docs/SECURITY.md documents the threat
model and what the connector can and cannot guarantee. 37 new tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs
(cherry picked from commit ee59322)
Socrata occasionally migrates a portal's domain (e.g. data.sfgov.org ->
data.sf.gov) and 301s every path on the old one. httpx.AsyncClient
defaults to not following redirects, so a portal_url that lags a
rename got back a raw 301 with an HTML body: raise_for_status() didn't
catch it (only 4xx/5xx), and response.json() failed trying to parse
the redirect page. That broke get_schema/get_dataset/query_dataset/
execute_sql while search_datasets kept working, since the Discovery
API treats old/new Socrata domains as aliases for search.

Hit this live: data.sfgov.org started 301-redirecting to data.sf.gov,
and every SODA3 tool failed until the deployment's portal_url was
updated and this fix landed.

Adds follow_redirects=True to the SODA client (the Discovery client is
unaffected — api.us.socrata.com is stable infra, not a per-portal
domain) plus a test asserting it's set.

Claude-Session: https://claude.ai/code/session_01HuLvMhEAJgh6M1Rih5gWpD

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 1f3b400)
…s everywhere (#20)

Follow-up to #19. #19 made the Socrata SODA client follow redirects so a
renamed portal domain (data.sfgov.org -> data.sf.gov) keeps working, but
httpx only strips Authorization on cross-origin redirects — not custom
headers such as Socrata's X-App-Token. A lapsed portal domain can be
re-registered by someone else, who would then receive the token (and, for
the other providers, any api key) from every deployment whose portal_url
lags the rename.

core/base_plugin.py: _create_http_client gains protect_headers= and
trusted_hosts=. When set, it follows redirects and attaches a request event
hook that drops those header names on any hop (initial or redirect) whose
host is not the configured portal/base host, a subdomain of it, or a listed
extra host. Shared _host_is_trusted / _trusted_request_hosts helpers.

All four plugins opt in with their credential header:
- Socrata SODA client: X-App-Token (replaces the bare follow_redirects=True)
- CKAN, Opendatasoft: Authorization (also newly follow redirects)
- ArcGIS hub + feature clients: Authorization, trusting *.arcgis.com and
  trusted_service_hosts (where feature services live)

A legitimate rename is indistinguishable from a hijack, so the credential is
not forwarded to the new host; the request still follows through
unauthenticated (Socrata's app token is only a rate-limit key; public data on
the others still resolves). Operators should update portal_url to the new
domain.

Tests: base hook coverage (same/subdomain kept, untrusted dropped + logged,
rename drops, extra trusted host retained, follow_redirects forced) and a
Socrata assertion that the SODA client opts in. docs/SECURITY.md documents
the scoping. 624 tests pass; no new ruff findings vs main.

Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 260fbc0)
…#14)

* feat(ckan): surface catalog metadata and add list_datasets + get_catalog_stats

A Boston eval comparing an agent with and without OpenContext flagged the
connector's top gap: get_dataset summarized package_show into prose and
dropped metadata_modified, license, organization, tags, and every resource's
url/created/last_modified (the resource URL alone would have dated the
year-series resources the agent had to leave null), and there was no way to
count the catalog, so catalog-wide statistics were asserted instead of
counted.

core/base_plugin.py (shared by all plugins):
- short_date (ISO / epoch s / epoch ms -> YYYY-MM-DD), human_size,
  format_search_header (catalog-wide total + "showing a-b"), and
  display_portal_url: a portal-supplied URL is echoed only when its host is
  the portal/API host or a subdomain (or an extra trusted host); otherwise
  only "(external: hostname)" is shown. Generalizes ArcGIS _display_url.

plugins/ckan/plugin.py:
- get_dataset now prints organization (title + slug), license (title + id),
  created/modified, tags, groups, and per resource: created, modified
  (last_modified or metadata_modified), size, DataStore flag, gated download
  URL and description. New max_resources arg (default 50, max 500) with an
  explicit "... and N more" line.
- search_datasets header uses package_search's count; rows show
  organization, modified date, resource count + formats, tag count.
- New list_datasets: exact-match filters (organization/tag/format/license/
  group), sort enum (default metadata_modified desc), limit/offset, total
  count. Filters are whitelisted and values quoted as escaped Solr phrases
  (_build_fq), so model input cannot alter the fq query. fl is deliberately
  not used (it collapses organization and drops resources).
- New get_catalog_stats: package_search rows=0 + facet.field for
  organization/tags/res_format/license_id/groups, optional query/filters,
  values sorted by count; falls back to legacy `facets` on older CKAN.
- Guidance strings point the model between the three catalog tools.

Tests: 22 new (Solr escaping/whitelist, enriched formatters incl. external
and non-http resource URLs, max_resources clamp/truncation, count header,
list_datasets request shape and invalid sort, catalog stats request shape,
sorting, legacy fallback, empty facets, unknown facet). Docs updated
(BUILT_IN_PLUGINS, ARCHITECTURE incl. previously missing aggregate_data,
SECURITY).

Verified live against data.boston.gov: 235 datasets, facet buckets,
boston-311-org listing sorted by modified, 311 dataset with 21 dated
resources.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

* style: datetime.UTC, ClassVar test fixtures, drop unused noqa

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

* style: ClassVar on remaining test fixtures

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

* style: hoist CKAN helper imports to module top (CI ruff E402)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 384b3d0)
…ta, ODS, ArcGIS (#16)

Companion to the CKAN catalog work (#14): the same shape of loss existed in
the other three providers — none passed the search total to its formatter
(Discovery resultSetSize / ODS total_count / Hub numberMatched), and each
get_dataset dropped dates, license, publisher/attribution and counts the
API had already returned.

Socrata: _discovery_search returns the envelope; search header uses
resultSetSize and rows show modified date, column count, downloads, source.
get_dataset adds source/attribution (link host-gated), license name + id
(terms link host-gated) instead of a stringified dict, created/published/
metadata-modified/rows-updated dates, rows/columns/downloads/views/
provenance; empty values are omitted instead of "N/A".

Opendatasoft: _catalog_search returns the envelope; header uses total_count
(the pattern _tool_query_data already used) and rows show publisher,
modified, records. get_dataset adds publisher, license (+ gated URL),
attribution, data/metadata processed dates, field count, references.

ArcGIS: _search_hub returns results + numberMatched; rows show owner,
created/modified, record count. get_dataset adds organization, last-edit
date, size, record count, categories, type keywords, access information;
empty fields are omitted. _display_url now delegates to the shared
display_portal_url with *.arcgis.com + trusted_service_hosts, so untrusted
hosts render as "(external: host)" like every other provider.

Tests: base helper coverage (short_date, human_size, display_portal_url,
format_search_header) plus enrichment tests per provider. Docs updated.

Claude-Session: https://claude.ai/code/session_01U7kJqaZ5pxpAKbZBr1biBs

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 0ae1dcb)
Socrata's developer console issues two different credential types:
a single-string App Token (Developer Settings -> App Tokens), and an
API Key pair (Key ID + Key Secret, Developer Settings -> API Keys)
meant for HTTP Basic Auth on authenticated requests. This plugin only
implements the former, sending app_token bare as X-App-Token.

Pasting an API Key's Key ID in as app_token fails silently for some
tools and not others: search_datasets/get_dataset keep working
(catalog/metadata calls tolerate it), but query_dataset fails with
"Invalid app_token specified" (403) on /resource/{id}.json — the
actual SODA3 data-query endpoint. Confirmed against data.ny.gov,
data.cityofnewyork.us, data.lacity.org, and data.sfgov.org while
deploying new portals.

Updates the config schema field description/validation error and
BUILT_IN_PLUGINS.md to steer setup toward the right credential type.
No auth code changes needed — a bare App Token is sufficient since
Socrata open-data portals are all public.

Claude-Session: https://claude.ai/code/session_01HuLvMhEAJgh6M1Rih5gWpD

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit e1ff187)
…CP adapter, test layout

Adopts thealphacubicle/OpenContext develop as the fork's baseline, carrying
every fork improvement already ported there (shared base layer, Opendatasoft,
portal-content guardrails, redirect credential scoping, catalog metadata).
Brings in the Typer CLI (opencontext configure/deploy/serve/...), the GCP
Cloud Functions adapter and terraform, the unit/security/integration/smoke
test layout with the upstream suites, and upstream docs and AI context.
Drops the Boston-specific deploy script and tfvars; deployment is now
driven by the CLI from a generated config.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
- sqlparse>=0.6.0, aiohttp>=3.14.3, pytest>=9.0.3 (Dependabot bumps the fork
  had already taken); uv.lock re-resolved.
- CI test job runs on Python 3.11 and 3.12 with upstream's coverage gate.
- Drop the sync-develop workflow: this fork has no develop branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
- Local dev servers (`opencontext serve`, tests' local_server) no longer echo
  exception text to the client; details stay in the log
  (py/stack-trace-exposure).
- Tests compare URL sets / use anchored regexes instead of substring checks
  CodeQL reads as URL sanitization (py/incomplete-url-substring-sanitization).
- Workflows declare `permissions: contents: read` at the top level
  (actions/missing-workflow-permissions); release.yml jobs keep their own
  write grants.
- Drop notify-bulletin.yml: it is gated to thealphacubicle/OpenContext and
  posts to their bulletin endpoint, so it never runs here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX
Adopt thealphacubicle develop as baseline: CLI, GCP adapter, test layout
… f=json, schema fallback, optional app token (#26)

* fix(arcgis): auto-trust Hub-referenced service hosts; schema fallback; wizard prompt

Findings from the 51-jurisdiction open-data audit (Open_Data_Inventory,
reports/gaps_and_improvements.md §4.5, §4.10) and thealphacubicle PR thealphacubicle#75:

- Hub catalogs routinely reference Feature Services on city GIS domains
  (gis.charlottenc.gov, maps2.dcgis.dc.gov, gis.indy.gov). The SSRF guard
  refused them unless each host was pre-listed, so Charlotte, Denver and
  Indianapolis lost query_data/get_schema entirely. New
  `auto_trust_hub_services` (default true) accepts a Hub-referenced URL when
  it is https, on a public DNS name (IP literals, single labels and
  .internal/.local names are always refused), and has an ArcGIS REST
  service path. The bearer token is never sent to auto-trusted hosts
  (protect_headers already scopes it to portal/arcgis.com/allow-list).
  Explicit trusted_service_hosts still wins, and operators can turn
  auto-trust off.
- Refusals are structured: the error starts with
  `untrusted_service_host: '<host>'` and names the config key, instead of
  surfacing as a generic "not trusted"/JSON-decode failure that three audits
  misdiagnosed.
- get_schema derives the field list from a one-row /query when the layer
  metadata endpoint returns non-JSON or an ArcGIS error envelope (Columbus,
  Indianapolis: metadata endpoint flaky while /query worked).
- Service URLs are whitespace-stripped before validation.
- `opencontext configure` prompts for trusted_service_hosts in the ArcGIS
  step (ported from PR thealphacubicle#75 by cstirry, mapped onto this field name).

Docs: BUILT_IN_PLUGINS.md, SECURITY.md, ARCHITECTURE.md, config-example.yaml.
Tests: auto-trust acceptance/refusal matrix, plugin-level on/off, schema
fallback paths, CLI prompt parsing. 1105 passed, 94% coverage, ruff clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX

* fix(plugins): port #23 — f=json as a param, first-layer resolution, optional Socrata token

Ports #23 (fix/connector-gaps-cdc-hud, found against
data.cdc.gov and HUD's ArcGIS Hub, 2026-09-20/21) onto the new layout, merged
with the auto-trust and schema-fallback work on this branch:

- ArcGIS get_schema built `{url}/0?f=json` and passed `params={}`; httpx
  replaces the URL's query string with params, so the request went out bare,
  ArcGIS answered HTML, and `.json()` failed on every dataset. This is the
  root cause behind the "JSON-decode" schema failures the audit logged for
  Columbus and Denver. The format now travels as `params={"f": "json"}`; the
  one-row-query fallback stays for services whose metadata endpoint is
  genuinely broken, and reports a ValueError naming both endpoints when the
  query path is unusable too.
- `_resolve_layer_url` reads the service description once, uses the first
  layer id (then the first table, then 0), caches per service, and keeps an
  explicit `/N`. Services whose only layer is not id 0 (HUD Low-Mod Income by
  Tract = layer 4, Opportunity Zones = layer 13) failed with "Invalid URL".
  get_schema and _query_features both use it.
- Socrata `app_token` is optional: blank means no X-App-Token header.
  data.cdc.gov serves SODA3 untokened but rejects an invalid token (403), so
  a deployer without a token could not run the plugin at all (audit §4.1).
  Docs updated.

Tests from #23 land under tests/unit/plugins/{arcgis,socrata}/.
1111 passed, 94% coverage, ruff clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX

* ci: run the Terraform check on every PR so the required status always reports

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TXm8YpvmTqNHgpRuoxtpoX

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Bumps [anyio](https://github.com/agronholm/anyio) from 4.12.0 to 4.14.2.
- [Release notes](https://github.com/agronholm/anyio/releases)
- [Commits](agronholm/anyio@4.12.0...4.14.2)

---
updated-dependencies:
- dependency-name: anyio
  dependency-version: 4.14.2
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [starlette](https://github.com/Kludex/starlette) from 0.52.1 to 1.3.1.
- [Release notes](https://github.com/Kludex/starlette/releases)
- [Changelog](https://github.com/Kludex/starlette/blob/main/docs/release-notes.md)
- [Commits](Kludex/starlette@0.52.1...1.3.1)

---
updated-dependencies:
- dependency-name: starlette
  dependency-version: 1.3.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
format_records defaulted to max_display=10, so query_data on ArcGIS (and
the CKAN/Socrata SQL and aggregate paths, which hard-coded 10, and CKAN
query_data, which hard-coded 5) fetched up to `limit` rows but showed
only the first few. Clients re-paged for rows they had already received.

max_display now defaults to None (all records). A character budget just
under the framing cap drops whole records from the end and says
"Showing N of M record(s)" instead of cutting a record mid-value.

Closes #28

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Resolves the open Dependabot alerts on uv.lock and one advisory pip-audit
reports that Dependabot has not raised yet:

- urllib3 2.6.3 -> 2.8.0 (alerts #16, #17, high)
- msgpack 1.1.2 -> 1.2.2 (alert #30, high)
- idna 3.11 -> 3.20 (alert #18)
- pip 26.0.1 -> 26.2.1 (alerts #14, #15, #31, #39; dev-only, via pip-audit)
- click 8.3.1 -> 8.3.3 (PYSEC-2026-2132; 8.5 deprecates APIs typer uses)

Lockfile only; no pyproject or requirements.txt changes.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
query_data gains `offset` (resultOffset), `order_by` (orderByFields,
validated to field names with optional ASC/DESC) and `format`
(text | json | csv). Replies now start with "Returned N of M matching
record(s) (offset, limit)" and end with "Next page: offset=..." when more
records match. The total comes from a returnCountOnly request, skipped
when a short first page already proves it; a failed count falls back to
exceededTransferLimit.

The shared base gains render_rows(), which renders text blocks, a JSON
array (one object per line) or CSV within the response budget, dropping
whole rows and reporting how many were shown so the next offset is
exact. format_records is now a wrapper over it.

ArcGIS plugin_version 1.0.0 -> 1.1.0.

Closes #29

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

A Feature Service can hold several layers and tables, but get_schema and
query_data only reached the default one. HUD's Small Area FMRs keep the
rents in table 1 behind a geometry-only layer 0, so they were unreachable.

- get_schema and query_data take an optional integer `layer`. The id is
  checked against the service's layers and tables; an unknown id is
  refused with the list of valid ones.
- get_dataset lists the service's layers and tables (id, name, geometry
  or "table") and marks the default. The lookup is skipped for untrusted
  hosts and never fails the call.
- The service description is read once per service and cached; the
  default-layer resolution now reuses it.

Closes #30

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Totals by area needed every row pulled and summed client-side: five-county
voucher sums took 23 paged query_data calls. aggregate_data sends one
outStatistics query with groupByFieldsForStatistics instead.

- statistics: [{type: count|sum|avg|min|max|stddev, field, as}]; a count
  without a field counts records by the object-id field.
- group_by, where, having, order_by, limit, layer, format (text|json|csv).
- Field and group names must be plain identifiers and exist in the layer
  schema (case-insensitive; sent with the schema's spelling). where and
  having use WhereValidator; order_by uses validate_order_by. All input
  checks run before any request.
- Layers reporting supportsStatistics: false are refused.
- The guidance line says nulls are excluded, so sums over suppressed
  values are lower bounds.

get_schema, query_data and aggregate_data now share _layer_url_for
(item -> trusted layer URL) and _layer_fields (fields + metadata, with
the one-row-query fallback).

Closes #31

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…38)

Descriptions were cut at 300 characters of raw HTML, hiding definitions
and suppression rules; the only date shown was the Hub item's; answers
carried nothing to cite.

- core/portal_content.html_to_text (stdlib html.parser): block tags to
  lines, list bullets, entities decoded, script/style dropped. Applied to
  item descriptions, snippets, licence and access information.
- get_dataset shows the whole description (capped at 12,000 characters
  with a truncation notice); search keeps a 300-character excerpt.
- get_dataset adds the default layer's data and schema edit dates from
  editingInfo (HUD FMR: 2025-09-30). Layer descriptions are cached.
- get_schema adds string field lengths and coded-value/range domains.
- query_data and aggregate_data end with a source block: dataset title
  and Hub page, layer queried, data edit date (only from the cache, so no
  extra request), licence, retrieval time.

ArcGIS metadata XML (field labels) is left for a follow-up.

Closes #32

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- ARCHITECTURE.md: the ArcGIS tool table listed four tools with their old
  arguments; it now lists get_schema and aggregate_data, the layer, paging
  and format arguments, and links to BUILT_IN_PLUGINS.md.
- BUILT_IN_PLUGINS.md: search/aggregation tools take `query`, not `q`;
  replace the "appends /0" note with first-layer resolution and `layer`;
  schema fallback path uses the resolved layer; WHERE validation covers
  having and order_by; get_schema row mentions lengths and domains.
- CUSTOM_PLUGINS.md: document render_rows.
- SECURITY.md: ArcGIS order_by / aggregate_data / layer input checks,
  html_to_text, and render_rows in the structure-forgery row.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants