Skip to content

feat(arcgis): aggregate_data tool (outStatistics + groupBy) (#31) - #37

Merged
sgarcese merged 1 commit into
mainfrom
feat/arcgis-aggregate-data
Sep 25, 2026
Merged

sgarcese merged 1 commit into
mainfrom
feat/arcgis-aggregate-data

Conversation

@sgarcese

Copy link
Copy Markdown
Owner

Closes #31 (fork-side tracking for thealphacubicle#83). Mirrors CKAN's aggregate_data, and uses layer (#36) and the json/csv output (#35).

Problem

Totals by area needed every row pulled and summed client-side. Summing housing-choice vouchers for five counties meant paging through 222 tracts in 23 query_data calls, even though the layer supports server-side statistics. (get_aggregations counts facets in the catalog; it does not aggregate data.)

Change

New tool aggregate_data

  • statistics: [{"type": "sum", "field": "HCV_PUBLIC", "as": "vouchers"}, …]. The types are count, sum, avg, min, max and stddev; a count with no field counts records by the object-id field.
  • Also takes group_by, where, having, order_by, limit (groups), layer and format.
  • It sends one query with outStatistics and groupByFieldsForStatistics, plus havingClause and orderByFields when given.

Validation. All checks run before any request, and all of them are covered in tests/security:

  • The statistic type must be in the list above; fields and output names must be plain identifiers.
  • where and having go through WhereValidator, and order_by through validate_order_by.
  • Field and group names must exist in the layer's schema. They are matched case-insensitively and sent with the schema's spelling. An unknown name says so and points to get_schema.
  • Output names must be distinct from each other and from the group fields.
  • A layer that reports supportsStatistics: false is refused, with a pointer to query_data.

Output

  • "Aggregated N group(s):" in text, json or csv. A trimmed result says "showing the first L (limit)", and hitting the service's transfer limit is reported.
  • The guidance line after the data says that nulls are excluded, so a sum over suppressed values (HUD voucher tracts under 11) is a lower bound.

Refactor. get_schema, query_data and aggregate_data now share:

  • _layer_url_for: item → trusted layer URL, with an optional queryable-type check.
  • _layer_fields: the layer's fields plus its raw metadata, with the existing one-row-query fallback.

There's no behaviour change for the existing tools. query_data now validates where before the item lookup instead of after.

Other: docs/BUILT_IN_PLUGINS.md is updated.

Tests

  • New tests/unit/plugins/arcgis/test_aggregate_data.py (23 tests):
    • how the request is built, and case-insensitive field matching
    • count falling back to the object id, having/order_by, and layer selecting a table
    • text and csv output, the limit notice, the transfer-limit notice, and no groups
    • unknown fields and group fields, a layer without statistics support
    • duplicate and clashing output names, sum without a field
    • malformed statistics never reaching the service, and an unknown format
  • New tests/security/test_arcgis_aggregate_guards.py (24 tests): injection strings in the statistic field, output name, group_by, having and order_by are refused before get_dataset or any request.
  • pytest -n auto: 1219 passed. ruff check and ruff format are clean.

Live against hudgis-hud.opendata.arcgis.com

Housing Choice Vouchers by Tract (8d45c34f…), sum of HCV_PUBLIC and count of GEOID, grouped by STATE/COUNTY for the five counties, order_by: vouchers DESC. This took 3 requests (service description, layer metadata, one statistics query) instead of 23 paged calls:

vouchers,tracts,STATE,COUNTY
2609,82,18,141
767,45,18,039
683,53,26,021
207,30,18,091
53,12,18,099

These are the same figures as the issue. LIHTC grouped by COUNTY_LEVEL (sum of LI_UNITS, record count) also returns all five counties.

Note

WhereValidator blocks destructive keywords (DROP, DELETE, …) but not UNION/SELECT. That is true of query_data's where today, and having uses the same validator. ArcGIS standardized queries reject subqueries on the server, so this adds no new exposure, but tightening the validator could be a follow-up.

🤖 Generated with Claude Code

Totals by area needed every row pulled and summed client-side: five-county
voucher sums took 23 paged query_data calls. aggregate_data sends one
outStatistics query with groupByFieldsForStatistics instead.

- statistics: [{type: count|sum|avg|min|max|stddev, field, as}]; a count
  without a field counts records by the object-id field.
- group_by, where, having, order_by, limit, layer, format (text|json|csv).
- Field and group names must be plain identifiers and exist in the layer
  schema (case-insensitive; sent with the schema's spelling). where and
  having use WhereValidator; order_by uses validate_order_by. All input
  checks run before any request.
- Layers reporting supportsStatistics: false are refused.
- The guidance line says nulls are excluded, so sums over suppressed
  values are lower bounds.

get_schema, query_data and aggregate_data now share _layer_url_for
(item -> trusted layer URL) and _layer_fields (fields + metadata, with
the one-row-query fallback).

Closes #31

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sgarcese
sgarcese merged commit 1c6e787 into main Sep 25, 2026
9 checks passed
@sgarcese
sgarcese deleted the feat/arcgis-aggregate-data branch September 25, 2026 03:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ArcGIS: aggregate_data (outStatistics + groupBy) — fork-first for upstream #83

1 participant