Repository navigation
feat(arcgis): aggregate_data tool (outStatistics + groupBy) (#31) - #37
Merged
Merged
Conversation
Totals by area needed every row pulled and summed client-side: five-county
voucher sums took 23 paged query_data calls. aggregate_data sends one
outStatistics query with groupByFieldsForStatistics instead.
- statistics: [{type: count|sum|avg|min|max|stddev, field, as}]; a count
without a field counts records by the object-id field.
- group_by, where, having, order_by, limit, layer, format (text|json|csv).
- Field and group names must be plain identifiers and exist in the layer
schema (case-insensitive; sent with the schema's spelling). where and
having use WhereValidator; order_by uses validate_order_by. All input
checks run before any request.
- Layers reporting supportsStatistics: false are refused.
- The guidance line says nulls are excluded, so sums over suppressed
values are lower bounds.
get_schema, query_data and aggregate_data now share _layer_url_for
(item -> trusted layer URL) and _layer_fields (fields + metadata, with
the one-row-query fallback).
Closes #31
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #31 (fork-side tracking for thealphacubicle#83). Mirrors CKAN's
aggregate_data, and useslayer(#36) and the json/csv output (#35).Problem
Totals by area needed every row pulled and summed client-side. Summing housing-choice vouchers for five counties meant paging through 222 tracts in 23
query_datacalls, even though the layer supports server-side statistics. (get_aggregationscounts facets in the catalog; it does not aggregate data.)Change
New tool
aggregate_datastatistics:[{"type": "sum", "field": "HCV_PUBLIC", "as": "vouchers"}, …]. The types arecount,sum,avg,min,maxandstddev; acountwith no field counts records by the object-id field.group_by,where,having,order_by,limit(groups),layerandformat.outStatisticsandgroupByFieldsForStatistics, plushavingClauseandorderByFieldswhen given.Validation. All checks run before any request, and all of them are covered in
tests/security:whereandhavinggo throughWhereValidator, andorder_bythroughvalidate_order_by.get_schema.supportsStatistics: falseis refused, with a pointer toquery_data.Output
Refactor.
get_schema,query_dataandaggregate_datanow share:_layer_url_for: item → trusted layer URL, with an optional queryable-type check._layer_fields: the layer's fields plus its raw metadata, with the existing one-row-query fallback.There's no behaviour change for the existing tools.
query_datanow validateswherebefore the item lookup instead of after.Other:
docs/BUILT_IN_PLUGINS.mdis updated.Tests
tests/unit/plugins/arcgis/test_aggregate_data.py(23 tests):countfalling back to the object id,having/order_by, andlayerselecting a tablelimitnotice, the transfer-limit notice, and no groupssumwithout a fieldstatisticsnever reaching the service, and an unknownformattests/security/test_arcgis_aggregate_guards.py(24 tests): injection strings in the statistic field, output name,group_by,havingandorder_byare refused beforeget_datasetor any request.pytest -n auto: 1219 passed.ruff checkandruff formatare clean.Live against
hudgis-hud.opendata.arcgis.comHousing Choice Vouchers by Tract (
8d45c34f…), sum ofHCV_PUBLICand count ofGEOID, grouped bySTATE/COUNTYfor the five counties,order_by: vouchers DESC. This took 3 requests (service description, layer metadata, one statistics query) instead of 23 paged calls:These are the same figures as the issue. LIHTC grouped by
COUNTY_LEVEL(sum ofLI_UNITS, record count) also returns all five counties.Note
WhereValidatorblocks destructive keywords (DROP, DELETE, …) but notUNION/SELECT. That is true ofquery_data'swheretoday, andhavinguses the same validator. ArcGIS standardized queries reject subqueries on the server, so this adds no new exposure, but tightening the validator could be a follow-up.🤖 Generated with Claude Code