Skip to content
isglobal-brgePublic

About

To be supplied

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

347 Commits

Folders and files

Repository files navigation

dsFlower

dsFlower is the node-side DataSHIELD package for running Flower federated learning under a privacy policy controlled by the data custodian. It is installed on each Opal/Rock node and pairs with the researcher-side dsFlowerClient package.

The node, its custodial root secret and its installed canonical runner are trusted. The researcher, submitted configuration, Flower coordinator and uploaded code are not. The supported privacy contract is therefore deliberately narrow: every numeric model update released by a SuperNode is produced by the node-installed, hash-pinned runner under a server-owned per-training (epsilon, delta) contract. dsFlower is not an unrestricted remote Python executor.

Package roles

Package Installed at Responsibility
dsFlower Data-owning Opal/Rock node Validate and stage local data, bind the per-training privacy contract, enforce the selected DP mechanism, run a Flower SuperNode and minimize DataSHIELD egress.
dsFlowerClient Researcher workstation Create declarative requests, verify runner compatibility, operate the SuperLink and coordinate cleanup.

Raw rows, images, masks and patient identifiers remain on the data node. Model updates travel through Flower only after the node-side mechanism has processed them. Each node acts as a trusted curator and applies central DP to its local dataset before egress; this is not formal local DP (LDP). dsFlower does not implement Secure Aggregation, so an untrusted/public coordinator is rejected unless the custodian explicitly opts in with dsflower.allow_untrusted_coordinator = TRUE.

Installation

remotes::install_github("isglobal-brge/dsFlower")

The configure script prepares two node-owned Python runtime families: a PyTorch/Opacus environment for neural and vision training, and a small native-tree environment for trusted tree training, validation and executable capability probes. The latter pins Flower 1.31.0, NumPy 2.4.6, pandas 3.0.3, PyArrow 23.0.1 and cryptography 46.0.7 exactly. It does not install upstream XGBoost, LightGBM or CatBoost: XGBoost remains in its separately verified native bundle, while the other two names identify dsFlower-style numeric engines.

Configure the curated XGBoost bundle

dsflower.xgboost_bundle_root is a node/custodian configuration option: the absolute path to a platform-specific directory containing the verified manifest.json and the XGBoost and dsFlower DP primitive shared libraries under lib/. It is not an analyst request parameter or a path to an upstream XGBoost installation. The ordinary R package installer does not build this bundle.

From a dsFlower source checkout, a custodian with Git, Python 3, Rust/Cargo 1.88 or newer, CMake 3.18 or newer, and a C/C++17 toolchain can build it as follows (the output directory must not already exist):

work_dir="$(mktemp -d)"
native/xgboost/scripts/fetch_upstream.sh "$work_dir/xgboost"
native/xgboost/scripts/apply_patches.sh "$work_dir/xgboost"
native/xgboost/scripts/verify_patched.sh "$work_dir/xgboost"
native/xgboost/scripts/build_bundle.sh "$work_dir/xgboost" /srv/dsflower/xgboost-bundle
python3 native/xgboost/scripts/verify_bundle.py /srv/dsflower/xgboost-bundle

Create the destination's parent directory beforehand. Deploy the complete bundle in a canonical absolute path without symlinks. On POSIX, its files, directories and parent chain must be owned by root or the node account and must not be group or world writable. Ensure the node service account can read the files and traverse the directories; a bundle built by another custodian account needs an explicit ownership or read-permission handoff. Configure equivalent owner-only write permissions on Windows. The native libraries must load without LD_* or DYLD_* loader overrides. The verifier rejects extra files, so store associated licenses and operational logs outside the bundle directory. See the native build documentation for source pins, platform suffixes and the complete native test.

Configure each DataSHIELD node's R process before serving analyst requests:

options(dsflower.xgboost_bundle_root = "/srv/dsflower/xgboost-bundle")

Alternatively set DSFLOWER_XGBOOST_BUNDLE_ROOT in the node service environment; a nonempty environment value takes precedence over the R option (which also supports the default.dsflower.xgboost_bundle_root option fallback). The native-tree Python runtime must also be provisioned; its default location is native-tree under the node's configured venv root, with an optional dsflower.native_tree_runtime / DSFLOWER_NATIVE_TREE_RUNTIME override.

Each capability check verifies the bundle and runs a synthetic public training, sanitization, ensemble and prediction probe. An absent, mismatched, tampered or unloadable bundle, or a failed executable probe, keeps XGBoost unavailable and training fails closed. There is no fallback to upstream or non-private XGBoost. Analysts can submit ds.flower.model.xgboost() through the normal R API only when all selected nodes advertise the resulting capability.

Computation contracts

Request Enforced node-side behavior
Declarative neural/vision specification Opacus DP-SGD with per-example or, when a server-selected patient identifier exists, per-patient clipping and noise.
HookApp Complete-update clipping and conservatively RDP-calibrated Gaussian output perturbation; optional fixed-block sample-and-aggregate only inside the required sandbox.
Private model validation One fixed Gaussian-noised vector of bounded per-unit sufficient statistics; only pooled metrics are post-processed by the ServerApp.

Declarative specifications are data, not researcher code. They provide the granularity of an nn.Module because the trusted runner owns the training loop and can observe per-sample gradients. Adding a reviewed operation to this declarative vocabulary is the safe extension path.

The survival contracts pytorch_aft and pytorch_discrete_hazard require custodian-configured patient privacy and an explicit stable subject identifier. They accept numeric baseline features and ordered (time, event) targets in days from baseline, with public resolution and administrative horizon. AFT supports Weibull or log-normal likelihoods with fixed public dispersion in {0.5, 1, 2}. Discrete hazard uses a public grid of at most 64 intervals and one fixed-width outcome/mask vector per subject; its masked BCE sum is divided by public K. Both sample, clip and account once per subject, including invalid subjects. Duplicate rows, unusable identifiers and invalid outcomes become safe zero-loss subject contributions. Source rows and subjects are never removed or disclosed. Cox remains excluded because its risk-set loss couples different subjects.

Survival private validation, atomic holdout and CV release observed-status Brier at public horizons and bounded fitted NLL. Concordance remains a public-split or authorized analyst-local metric; private HPO remains unsupported. See private validation and CV for the layouts and calls. Predictive utility is an empirical question; implementing the contract does not establish a useful concordance score for any dataset or privacy budget.

A HookApp is more restricted than a general Flower App. It exposes initial_arrays() and local_update() and is never imported into the trusted parent. Arbitrary code cannot generically receive DP-SGD-level guarantees: static inspection cannot establish per-sample gradient independence or exclude side channels. A HookApp executes only when all of the following are true:

  • the custodian enables it;
  • a Bubblewrap filesystem/network boundary is available and explicitly attested;
  • the configured minimum-duration timing envelope is valid;
  • the uploaded package passes archive validation, scanning and hash pinning.

Both child functions receive only a bounded, hash-pinned public app_params object plus trusted round_index, num_rounds, task and class-count fields. Privacy, path, dependency, secret and runtime keys are reserved. Array count and shape are fixed for the run, and the Gaussian scale composes the complete round transcript. If a public execution gate is absent, the HookApp is not executed and the run is reported as available=false without a model artifact. A child crash, timeout or malformed result inside the sandbox becomes a zero delta and still receives the same Gaussian output mechanism. A failure in the trusted runtime after private execution starts is also reported only as unavailable, without its cause or fallback being accepted as a trained model.

The Hook timing envelope is defense in depth, not a formal constant-time guarantee: cleanup, process availability and storage behavior remain outside the numeric DP proof and require deployment-level quotas/isolation when in scope.

The dedicated native-tree ABI implements reviewed node-owned mechanisms for XGBoost, ExtraTrees, adaptive Random Forest, dsFlower LightGBM-style boosting and dsFlower CatBoost-style boosting. Every request pins public bounds, public cuts and a fixed typed parameter profile before private data are opened; the node owns privacy calibration and sticky custodial randomness. LightGBM-style and CatBoost-style are safe dsFlower numeric engines, not wrappers around the upstream binaries or model formats. Capability fields report fresh executable probes on each node; the engine list describes request syntax and is not a privacy permission catalogue.

The Random Forest mechanism uses a disjoint node-owned partition: each effective privacy unit contributes to exactly one tree. It is not upstream bootstrap/bagging Random Forest, and small cohorts can consequently have fewer effective units per tree. The public defaults (8 depth-4 binary trees; 4 depth-4 regression trees) are benchmark-oriented starting points, not private-size admission rules.

The four pure dsFlower engines are operational after the standard dependency-light runtime is provisioned. Native XGBoost remains fail-closed on a clean install: a custodian must separately build, verify and configure the platform-specific curated bundle. The installer does not compile or download that native trust artifact implicitly.

Private validation loads an already public declarative, native dsFlower vision, or sanitized native-tree model before opening the staged validation data. Each row or configured patient contributes one bounded histogram/sufficient-statistic vector. The node releases its sum once through the Gaussian mechanism; exact predictions, labels, counts and per-node metrics never leave the node. The ServerApp requires every selected node and derives binary, multiclass, ordinal, multilabel, bounded-regression, count and the patient segmentation/survival metrics only by post-processing the pooled DP vector. This is external validation when the assigned dataset is independent and resubstitution validation otherwise; it does not relabel reuse as cross-validation. If any expected node does not provide the fixed private release, the pooled artifact reports available=false and omits metrics. It never substitutes exact or zero-filled metrics, and this does not introduce a query-count lockout.

Atomic holdout is available for supported classification/regression/count tabular neural/native-tree training and native dsFlower 2D/3D vision models, patient segmentation and survival. Nodes derive the same secret-keyed row/patient split before training, spend the fixed 80/20 job budget on training and one pooled test release, and publish the model plus metrics only after the exact roster completes both phases. Vision paths and patient IDs are partitioned before pixel decode for ordinary vision. Segmentation uses its canonical patient assembly and excludes the opposite side from training and metric contributions.

Admitted radiomics tables

A complete radiomics data frame or Arrow table retrieved through dsImaging can be passed to ds.flower.fit() by its session symbol, including an unchanged Parquet round trip. This requires the coordinated dsImaging companion that registers exports against its private admitted patient roster. Admission requires the original export row order; reordered copies, changed values, subsets, duplicate/missing sample keys, unregistered generic dsHPC tables and revoked sources fail closed. Patient identity comes from the protected roster, never from caller-added attributes. dsFlower canonicalizes privacy units after admission for row-order-independent training. The companion prerequisite also applies when the public container has no patient-ID column.

Per-training privacy

Privacy is server-authoritative. The client cannot set epsilon, delta, clipping or HookApp controls. The custodian pins one positive epsilon/delta pair for each training, and its accountant composes that contract across the training's own rounds. There is no historical privacy-budget database, query quota or resource-specific privacy balance. Neighbourhood-store and gated-Hook cache capacities are separate storage admission limits. Fresh trainings compose under the usual conditional DP assumptions; repeated or nearby inputs can instead replay a neighbourhood anchor. Metric and threshold selection over one released DP model is ordinary post-processing; training a different model is a new per-training release.

The guarantee is accounted independently at each node. If one person can occur at m observed nodes, their federation-wide guarantee composes across those nodes (at most the sums of their epsilons and deltas); only genuinely disjoint node populations get parallel composition. A deployment needing one bound for overlapping sites must add a shared, person-level federation accountant.

The formal adjacency is bounded/replace-one with a fixed number of privacy units: neighbouring datasets replace one row, or all records belonging to one configured patient. This protects the values contributed by that unit; it is not an unbounded add/remove membership guarantee for a changing number of privacy units.

Deterministic, semantic-scoped randomness

Determinism prevents averaging only for an exact replay. Reusing one fixed noise vector for distinct or adaptively related queries is unsafe because correlated answers can cancel it. dsFlower derives deterministic randomness from a canonical, mechanism-bound semantic identity:

R = SHA256(frame("request-v3", canonical_public_request))
B = SHA256(frame("data-binding-v1", canonical_private_binding))
release_key = HMAC-SHA256(noise_root, frame("dsflower/semantic-prf/v3", R)
                                     || frame("data-binding", B))
subkey      = HMAC-SHA256(release_key, mechanism_axis)

The node secret is 32 bytes from the operating-system CSPRNG, created at runtime, stored outside staging with mode 0600 and never exposed to submitted code. Its file must be owned by the Rock service UID; its real parent directory may be owned by that UID or root, but must not be writable by group or other users. The parent is checked before and after the key is opened. When its path is explicitly provided as a process environment variable, the supplied Rock image bootstraps after runtime mounts are ready and before opening its port. Otherwise it deliberately waits for the first flowerInitDS(), when DataSHIELD profile options exist. Neither configure nor .onLoad() creates privacy state, and both Docker builds assert that no seed entered the image. A bootstrap storage error does not take Rock down; every private entry point retries and remains fail-closed until the mount is repaired.

A missing key is provisioned from OS entropy only before neighbourhood state is established. A retained local UUID pin, store, initialization lock or external UUID requires the original key. Existing malformed, wrong-owner or wrong-mode keys fail closed and are never silently replaced. Symlinks and unsafe parents remain rejected. Persist the secret, neighbourhood store and Hook cache across service/container replacements; changing the key creates an additional release domain, not a free retry.

The v3 public request R binds effective selected column/asset roles, canonical model specification, initial and incoming model contents, mechanism, raw privacy policy, local training/strategy semantics, round/fold and measured runtime. The private B binds complete selected source units as well as final tensors and private execution geometry. Distinct source data remain distinct even if they pool or bin to equal statistics. Canonical keyed row/patient ordering precedes computation. Symbols, handles, paths, row order, session/run/message IDs and pure server aggregation settings do not create noise axes. Existing authorization checks remain in force and private digests never leave the node.

All default neural server initializers are seeded from the canonical public model specification in an isolated Python/NumPy/Torch context; every CV fold shares the same start. Nodes retain array admission and content binding without recomputing round-one default arrays, preserving heterogeneous runtime support. Hook server initial_arrays receives the same RNG isolation, but nondeterministic output creates a new release per run. Public checkpoints retain their independent verification. See the randomness contract.

Version 0.7.2 mitigates the equality oracle described in isglobal-brge/dsFlower#7 with immutable neighbourhood anchors. For each public request R (including its incoming model and round), the node scans all retained anchors and returns the complete stored payload of the oldest anchor at distance d < k. If none is eligible, it computes the unchanged v3 content-bound release and appends one new anchor. Near inputs never become anchors. Newer anchors cannot displace an older eligible anchor, so replay is stable without per-input bindings.

Distance counts unordered canonical privacy units with multiplicity: d = max(|M| - c, |N| - c), where c = sum(min(M[t], N[t])). One insertion, deletion or replacement counts once; in patient mode a patient's complete selected records form one unit. The default k is the node's nfilter.subset (or its default. option), otherwise 3, with a floor of 2. Custodians can set dsflower.neighbourhood_k (or default.dsflower.neighbourhood_k); the effective value is frozen per R. Analysts cannot change it or force refresh.

An analyst can no longer test a one-unit difference against an existing release by obtaining a fresh answer inside that anchor's neighbourhood. This is a mitigation, not transcript DP: the hard k-1/k boundary and boundaries between anchors still distinguish some one-unit neighbours. An input must be at least k from every anchor to receive a fresh release. Small updates can therefore return stale models or statistics; uncertainty intervals do not include this staleness. No hit/near/fresh status, distance or anchor identifier is returned. Timing and availability remain outside the guarantee.

The rule covers every neural DP-SGD round, gated Hook release, all five native tree engines, private validation, holdout, CV fold training and OOF output, and association. Each round/fold/evaluation has its own R. Completed federation trajectories replay when the same incoming public models and eligible anchor choices recur; this is not a universal neighbouring-world transcript claim. Calibration, sensitivities, accounting, FedProx and the fresh R/B/K identity remain unchanged. Conditional DP mechanisms compose under their existing assumptions; no lifetime privacy budget is introduced.

Within one Flower run, a bounded claim ledger in the private staging directory reserves every operation/fold/round coordinate atomically before private work. It is mirrored into NodeState, survives ClientApp process restarts while that run's staging remains, and prevents concurrent processes from claiming the same coordinate. A changed payload cannot reuse a claim. Declarative requests fail closed once their in-memory reply has advanced; gated Hooks can replay earlier committed rounds from the durable cache after rechecking semantic data identity. A changed cache key cannot reuse a committed coordinate. This per-run control is distinct from a cross-training privacy-budget ledger; separate authorized trainings still compose under the custodian's deployment policy.

The trusted built-in tracks request strict deterministic Torch kernels. HookApps receive deterministic Python, NumPy and Torch seeds, and their final noise key is also bound to the validated clipped update. Arbitrary native user code cannot be certified deterministic by a static scanner. A durable node-owned cache therefore makes exact retry apply to every admitted HookApp, deterministic or not. Entries remain pinned throughout active runs. The neighbourhood store separately retains complete releases without eviction and is consulted before Hook execution. Hook execution is disabled by default and remains the deliberately weaker, custodian-gated extension path.

The current Gaussian sampler is a hardened Box--Muller construction over IEEE-754 values. ChaCha20 makes its finite random choices unpredictable, but it does not turn floating-point output into an exact continuous Gaussian. The implemented guarantee is therefore the documented computational/practical DP contract, not a claim of a formally verified finite-precision mechanism. Moving to a verified discrete-Gaussian sampler would require a new mechanism/accountant ABI and a reviewed privacy-contract change, rather than a drop-in RNG change.

DataSHIELD does not define a portable, connector-level confidential seed for Opal or Armadillo. Current Opal releases inject a short datashield.seed R option derived from the Opal service secret and log the resulting integer; Armadillo stores an administrator-editable nine-digit value in each profile. dsBase::setSeedDS() separately replaces R's mutable .Random.seed from an analyst-supplied integer and returns the resulting state. These are reproducibility/masking facilities, not confidential 256-bit DP mechanism keys. dsFlower intentionally neither depends on nor exposes them; the dedicated node secret is the fail-closed source of keyed release randomness.

Data and output minimization

Exact feature counts, sums and sums of squares are disabled. For tabular utility, the analyst may provide data-independent public lower/upper feature bounds; the same clipping and affine transform is applied during training and prediction. Without bounds, neural inputs remain unscaled but are locally coerced and saturated to [-1e6, 1e6].

Target preprocessing is also public and per-record. Classification strings or factors use an ordered target_levels; numeric labels may instead arrive already coded in [0, K-1]. Missing/unknown classification values map to the public code zero. Regression/count models require finite public target_bounds; values are coerced, non-finite/unparseable values map to the public bounds midpoint, and all values are clipped locally. Selected tabular features are likewise coerced to numeric and non-finite/unparseable values use the public bounds midpoint (or zero when bounds are absent); public numeric values/bounds are capped at 1e6 to stay below unsafe float32 center/span arithmetic. These are fixed per-record maps: no node derives a vocabulary, imputation value or range from its cohort, and no row is dropped.

Vision inputs are decoded under fixed node-side limits (256 MiB source and decoded payload, 32M elements) and embedded from paths in batches capped at 128 MiB; the full resized-image cohort is never retained. Raster, NIfTI, inline NRRD, MHA and DICOM inputs are supported. Detached NRRD payloads and .mhd sidecars are conservatively mapped to the same zero-image record as a corrupt input, before any sidecar is opened.

The declarative neural runner also totalises its arithmetic: safe division, finite saturation after every graph operation, parameters/intermediates bounded to magnitude 1e6, and loss-aware heads (1e6 for direct MSE regression, 30 for logits and log-links). Per-sample gradients are made finite coordinate-wise before Opacus applies the server-owned L2 clip, so an overflowing backward pass cannot suppress the Gaussian noise. Neural learning rates must be in (0, 10]. The same preprocessing and head saturation are replayed by local prediction helpers.

Run admission never inspects class or event frequencies. Such a check would turn prepare success/failure into a label-dependent oracle outside the DP mechanism. It does enforce the server-owned DataSHIELD minimum on the staged privacy-unit count: rows in row mode and distinct canonical patients in patient mode, including image runs. The threshold is at least nfilter.subset (default 3) and can be raised with dsflower.min_train_rows; a refusal returns one generic node error and never the exact count or shortfall. In patient mode unusable identifiers are collapsed into one fixed sentinel privacy unit; there is never a silent row-level fallback. This is deliberately conservative (it can protect several unidentified subjects together and reduce utility). A meaningful per-person interpretation still requires the custodian to provide a complete, stable identifier roster across releases.

flowerPrepareRunDS() validates and pins the complete per-training contract before it reads private table/file contents. flowerEnsureSuperNodeDS() checks the same pins again before launching the runner. Private-value preprocessing is totalised rather than returned as a success/error bit.

Node-side training logs and metrics are not returned through DataSHIELD, and Flower aggregation weights are fixed rather than revealing cohort size. The server exposes no log, metric or exact feature-statistics endpoint. The global model is the intended DP release; DP bounds an individual's influence on its distribution, not the possibility of every form of model inversion.

Binary 2D segmentation (experimental)

pytorch_resnet18_segmentation adds patient-level binary segmentation to the declarative neural track. This implementation is awaiting reviewer promotion; the preregistered BrEaST/BUS-BRA utility campaign has no F5 scores yet. Its presence in the API is not a validated practical-utility claim.

The public resnet18_layer2_128_v1 profile uses the frozen pretrained ResNet18 encoder through layer2, followed by a small trainable convolutional decoder. Only the decoder is DP-trained and released. The encoder remains node-local in evaluation mode with immutable BatchNorm statistics. Its exact checkpoint and 128-by-128 preprocessing are pinned; there is no random-weight fallback.

Each subject contributes the lexicographically first canonical image ID. Masks attached to that image are unioned; conflicting image paths invalidate that subject. The loss is the per-image mixture of BCE and Dice (alpha 0.5, smoothing 1; alpha 1 is the BCE ablation). Image inputs are PNG/JPEG, masks are single-channel PNG with an explicit 0,1 or 0,255 vocabulary. Invalid pairs produce fixed safe tensors with zero validity and remain in the subject census. Empty masks require an actual empty PNG or an explicit empty annotation.

Direct metadata tables require custodian options dsflower.dp_unit="patient", dsflower.patient_column, dsflower.image_data_root and dsflower.mask_data_root. The public schema declares distinct image-ID, image-path and mask-path columns, plus an optional empty-annotation column. The mask-path column is the target. Authorized dsImaging image bundles can use their declared mask_root asset and retain their existing patient roster checks. Neither route accepts an analyst-supplied filesystem root.

The decoder defaults to decoder_init = "random". Two digest-bound public initialisation routes are available in 0.7.1:

  • client:<path-to-bundle> validates researcher-local public material and declares its provenance and digests. This route supports research iteration.
  • resource:<handle-symbol> uses a custodian-registered checkpoint resource admitted through datashield.assign.resource() and flowerCheckpointInitDS(). This route supports curated institutional material, large artifacts, and nodes that admit no analyst-supplied material. The coordinator independently obtains the same public weights through its local public_checkpoint_file argument.

The custodian sets dsflower.public_initialisation to analyst_or_resource (default), resource_only, or none, optionally overridden by dsflower.public_initialisation.pytorch_resnet18_segmentation. Analysts cannot change this policy. The former public:<id> selector, registry beside the node secret, installer and manifest allowlist no longer authorize initialisation.

Both routes use a complete verified bundle, including the frozen encoder, checkpoint, manifest, provenance, licence, protocol and audit evidence. The node verifies a protected snapshot before private staging; the trusted runner verifies it again before private access. Incoming training arrays retain shape, dtype and value admission and their actual content identity; nodes do not compare them to expected round-one tensors. Missing or altered bundle material fails closed. Status returns public identity, provenance and geometry; it never exports checkpoint bytes. Canonical content identity excludes resource names, session symbols, locations and archive packaging, so aliases and repacking do not create another noise draw. Run manifests and release records retain the origin and verified digests.

See public initialisation procedures for exact Opal, Armadillo and DSLite commands, bundle requirements and policy configuration. The BUSI reference records retain the original hashes and evidence. The three original decoder binaries are still absent; synthetic tests do not establish recovery of those checkpoints. Public initialisation does not change the DP accountant, clipping, sampler, training algorithm, release cache or v3 identity machinery. Segmentation remains experimental; the mirror's licence declaration retains its original qualification.

Private validation, holdout and CV release bounded pooled foreground Dice. Mean patient Dice, IoU and empty/foreground strata remain public or authorized local metrics; local both-empty Dice is 1. Private HPO remains unsupported. See private validation and CV for the distinction and public-initialisation calls. Local predictions use the canonical 128-by-128 grid.

Custodian options

Options use the dsflower.* prefix, with the standard default.dsflower.* DataSHIELD fallback.

The supplied Rock image performs a pre-service key bootstrap only when a deployment explicitly provides DSFLOWER_NODE_SECRET_FILE. Otherwise it waits for the first session so Opal/Armadillo profile R options retain their normal precedence. The environment variable takes precedence over the R option, so a stale option cannot select another key path. Epsilon and all other policy controls remain normal DataSHIELD profile options.

Option suffix Default Meaning
dp_per_training_epsilon 1 Administrator-pinned epsilon for every training release; maximum 10.
dp_per_training_delta 1e-6 Administrator-pinned delta for every training release; maximum 1e-3.
dp_unit row Adjacency unit: exactly row or patient.
patient_column unset Required explicit stable identifier when dp_unit="patient"; never auto-detected.
image_data_root unset Custodian image root for direct segmentation metadata; declared paths must remain within it.
mask_data_root unset Custodian PNG-mask root for direct segmentation metadata; declared paths must remain within it.
public_initialisation analyst_or_resource Custodian admission policy: analyst_or_resource, resource_only, or none; optional per-contract suffix.
checkpoint_cache_dir checkpoints under tools::R_user_dir("dsFlower", "data") Protected service-owned cache outside staging, Hook mounts and node-secret storage.
dp_clipping_norm 1 Server-owned clipping bound.
node_secret_path Unix: /var/lib/dsflower/privacy/noise_root; Windows: %LOCALAPPDATA%/dsflower/privacy/noise_root Runtime-generated 256-bit node key; DSFLOWER_NODE_SECRET_FILE takes precedence when a deployment selects a service or secret-manager path.
neighbourhood_k nfilter.subset, then default.nfilter.subset, then 3 Positive integer, floored at 2, frozen per R; explicit/default dsFlower option takes precedence.
neighbourhood_max_anchors 256 Maximum retained anchors per R; only would-be-fresh inputs are refused at the cap.
neighbourhood_store_bytes 68719476736 Permanent logical/page-allocation store cap (64 GiB); exact and near replays are not charged.
neighbourhood_state_dir <node-secret-path>.neighbourhood Persistent owner-only store outside staging and Hook mounts.
neighbourhood_store_id unset Optional external lowercase UUID pin; must match established state.
release_cache_dir release-cache beside the node secret Persistent gated-Hook release cache, outside staging and Hook mounts; requires service-owned 0700 directories and 0600 regular files, with no symlinks.
release_cache_bytes 1073741824 Administrator-pinned logical byte capacity for encoded gated-Hook releases, bookkeeping and active-run reservations. Only a would-be-fresh anchor reserves the complete public worst-case run before Hook execution; exact/near anchors bypass this admission.
app_spool_root /var/lib/dsflower/appstore Private, persistent, service-owned upload spool; ephemeral and symlink paths are rejected.
max_fab_bytes 52428800 Per-FAB compressed upload cap.
app_spool_max_bytes 1073741824 Global logical-byte cap across all uploaded FABs and unpacked apps.
app_spool_ttl_seconds 86400 Incomplete-upload retention; installed catalogue apps persist until explicit deletion. Locked operations and staging-referenced apps are skipped by GC.
tunnel_chunk_bytes 524288 Maximum decoded payload in one DSI tunnel exchange; constrained to 16--512 KiB and negotiated with the client. The upper bound stays below DSI's expression-parser limit; larger streams are carried as multiple exact chunks.
tunnel_spool_max_bytes 1073741824 Per-direction tunnel spool cap; at least eight chunks and at most 64 GiB. TCP backpressure applies when full.
tunnel_loss_tolerance 180 Seconds without a relay heartbeat before the node forwarder exits; constrained to 5--86400.
hook_enabled FALSE Allow HookApp execution, subject to every other gate.
hook_sandbox_attested FALSE Custodian attestation of the Bubblewrap boundary.
hook_resource_isolation_attested FALSE Custodian attestation of externally enforced Hook resource isolation.
dp_sample_aggregate FALSE Enable fixed-block HookApp sample-and-aggregate when every sandbox gate is present.
dp_sa_blocks 8 Public, fixed HookApp block count, constrained to [2, 64]; never derived from private cohort size.
dp_egress_timeout 900 Hook child timeout in seconds.
dp_egress_time_pad 0 One minimum-duration envelope for the complete Hook release; zero disables execution. It must be at least dp_egress_timeout + 5, or dp_sa_blocks * dp_egress_timeout + 5 when sequential sample-and-aggregate is enabled. It is timing defense in depth, not a formal constant-time proof.
dp_egress_memory_mb 8192 Hook child address-space limit in MiB (512 to 131072).
dp_egress_file_mb 1024 Maximum size of any Hook child output file in MiB (16 to 16384).
dp_egress_processes 128 Hook child process/thread limit (1 to 1024, where supported).
allow_untrusted_coordinator FALSE Permit a coordinator to observe already-private per-node updates.

The permanent node-local store defaults to <node-secret-path>.neighbourhood. First release initializes it automatically and pins its random UUID at <node-secret-path>.neighbourhood-id, beside the secret rather than inside the store. The store uses owner-only directories/files (0700/0600), keyed unit fingerprints, MAC-authenticated records and complete payload bytes; it never stores raw source records or noise seeds. Records are verified before decoding, and a per-R lock serializes selection and durable anchor commit before release. There is no eviction, expiry or per-attempt ticket.

The default limits are 256 anchors per R and 64 GiB of store capacity. Fresh commits must fit both logical retained bytes (payloads, fingerprints and record overhead) and SQLite allocated pages plus fixed state headroom. Provision extra physical disk space for transient SQLite journals and filesystem allocation slack. Only an input that would create a new anchor is refused at a limit, with one stable error; exact and near replays remain available. A refusal substitutes for a fresh release and reveals the same fact that the input is far from every anchor. The shared byte cap also exposes a weak cross-user aggregate signal about prior store growth. These resource settings are not privacy parameters. Increase capacity with all retained state intact; never delete anchors to make space.

On Windows, the key parent and store directory require private inheritable ACLs: the service identity and trusted system/administrator principals may access the state; other grants, including read access, are rejected. Reparse points and hard links fail closed. SQLite uses its Windows VFS for durable commits; the first UUID pin is flushed and published without replacement using write-through publication. POSIX nodes retain file/directory fsync and flock. The Windows adapter has portable contract tests; Windows service and crash validation remains outstanding.

Missing or corrupt established state, missing original keys, unsafe permissions or a mismatched UUID fail closed. The permanent <node-secret-path>.neighbourhood-id.lock also detects loss of both the store and local UUID pin. Optional dsflower.neighbourhood_store_id externally pins the UUID and detects loss of all local store markers; without it, losing the store, UUID pin and initialization lock together can resemble first use. Stop all workers before restoring a consistent backup of the secret, store, UUID pin, permanent locks and retained Hook cache. MACs do not detect rollback to an older complete valid snapshot: avoiding rollback remains a custodian/storage assumption. Runtime upgrades create new public request domains, so drain jobs and upgrade both packages together; retain old state and treat new domains as additional releases. Existing per-release DP does not prove the private anchor-selection transcript DP.

hook_resource_isolation_attested=TRUE is an operator assertion, not an in-process control. Set it only when the SuperNode and every inherited Hook child are confined by cgroup v2 memory, PID and CPU limits (memory.max, pids.max, cpu.max) and their writable temporary filesystem is a size-limited tmpfs or quota-enforced volume. Bubblewrap/RLIMIT alone do not satisfy this second gate. Without both attestations, the time envelope and hook_enabled, a HookApp remains a data-independent no-op.

Every admitted HookApp uses the durable release cache. Identical semantic requests, including nondeterministic applications, replay the exact first released arrays and constant metrics without re-executing the Hook. The v3 key binds effective private data, source/column selections, verified application contents, public model, policy and round; run tokens, paths and cache capacity do not reroll the release. Neighbourhood anchors decide first: nearby data reuse a whole retained release. A fresh-path change to data or selections misses the exact Hook cache but cannot authorize a second release at a committed coordinate. The noised-zero failure outcome is cached in the same way.

Configure cache location and capacity through administrator dsflower.* or default.dsflower.* options. Analyst run configuration, nested app_params and manifest overrides cannot set them. Keep the directory on persistent local storage, outside staging, the application spool and all Hook filesystem mounts. Unsafe permissions, ownership, symlinks and nonregular files fail closed. Cache files contain only exact noised releases and operational bookkeeping; master seeds and noise keys are never persisted there.

R admission validates and retains the Hook cache settings without reserving capacity. After canonicalization, neighbourhood anchors decide first. Only a would-be-fresh anchor reserves capacity from public model bounds and the round count, before the Hook child executes: 65 MiB per round, plus 4 KiB + 2 KiB per round of run bookkeeping and 64 KiB of shared bookkeeping. Exact and near replays bypass that reservation even when the inner Hook cache is full. Provision extra physical disk headroom for SQLite journals and filesystem overhead; this logical reservation does not isolate storage faults. Concurrent identical requests serialize. Each active run pins every entry it uses, and cleanup closes that run before releasing its pins. Only unpinned entries are evicted, oldest first. Interrupted or crashed runs retain their uncertain pins until authoritative cleanup; do not manually delete their cache state to free space. Increasing the administrator's capacity or completing cleanup can restore Hook cache admission. The exact Hook cache's own retention policy is unchanged; the outer neighbourhood store permanently retains released payloads and serves exact/near matches even after the inner entry is evicted. Closed-run tombstones retain their metadata charge to reject late messages; enough retained bookkeeping can also exhaust admission capacity. If a crash loses the staging receipt or R handle, an administrator must stop the run's workers and explicitly close its stored run fingerprint with the trusted release_cache.py close command. Absence of a process or elapsed time never automatically removes uncertain pins. Automatic cleanup first confirms worker shutdown and closes an inner-cache run only if it has a reservation. Closing a replay-only run with no reservation is then a no-op and consumes no cache capacity. The administrator's explicit release_cache.py close command retains its unconditional tombstone behavior to block late admission. Preserve the node secret, both stores and local UUID pin across service replacements. Declarative tracks do not use this cache. Cache availability and hit/miss timing remain outside the numeric DP guarantee, and the minimum-duration envelope is unchanged.

Upload admission and writes are serialized by a node-global lock, so the physical byte cap is atomic across R sessions. There is no catalogue-entry or call-count quota. Before each admitted chunk, lazy TTL collection removes only expired incomplete uploads and only when their per-upload lock can be acquired without waiting. Verified installed apps persist until explicit flowerAppDeleteDS(). Pinning also records the server-generated run token in the app spool so active bytes remain immutable for the complete run.

The DSI transport is capability-bound and all-or-nothing. It does not add encryption to DataSHIELD itself: production Opal/Armadillo frontends must enforce TLS and certificate validation. The official dsFlowerClient can inspect the URL and retained verification options of DSOpal connections. A recognized Armadillo connection with a valid https:// URL is accepted automatically for connector parity; its frontend or reverse proxy remains authoritative for certificate and hostname validation. Unknown or unidentifiable connectors require an explicit per-site operator attestation. Plaintext HTTP is rejected by default and can be enabled only for exact named sites through the client-side dsflower.dsi_allow_insecure_http exception; that requires an independently trusted network layer. A loopback tunnel is operator-authorized only while its exact registered forwarder is alive and its post-bind ready marker exists. Failed startup kills the child and removes its registry/spool state; the client tears down every attempted site if any site fails. Exchanges use a per-session lock, bounded negotiated chunks and bounded client buffers. Spools are reset on a new SuperNode connection; acknowledged prefixes are compacted atomically under the same exchange lock while external offsets remain absolute. This keeps long sessions bounded without racing the R exchange, and TCP backpressure applies whenever a configured cap is reached.

DSI 1.8 can represent a per-node failure as a named NULL. Every mutating dsFlower path therefore requires an explicit, correctly named ok = TRUE ACK; NULL is never delivery. Upload and tunnel ACKs also bind the generation, offset, length and content identity. After an ambiguous response, only the byte-identical in-flight chunk may be replayed against an idempotent store: its geometry is never reduced or extended until that ACK is resolved.

Example Unix node policy (Windows services should set an absolute DSFLOWER_NODE_SECRET_FILE or default.dsflower.node_secret_path explicitly):

options(
  default.dsflower.dp_per_training_epsilon = 1,
  default.dsflower.dp_per_training_delta = 1e-6,
  default.dsflower.dp_unit = "row",
  default.dsflower.node_secret_path = "/var/lib/dsflower/privacy/noise_root",
  default.dsflower.release_cache_dir = "/var/lib/dsflower/privacy/release-cache",
  default.dsflower.release_cache_bytes = 1024^3,
  default.dsflower.app_spool_root = "/var/lib/dsflower/appstore",
  default.dsflower.hook_enabled = FALSE
)

The epsilon/delta pair describes one training. Its rounds are calibrated as one mechanism; later trainings do not consume a stored balance. Metric and threshold selection over one released DP model is post-processing; HPO or CV that trains new models creates new per-training releases.

The Python privacy runtime is constrained to the audited compatibility families: Flower 1.31.0 exactly, Torch 2.x, Opacus 1.x and torchvision 0.x. The dedicated native-tree runtime uses the exact dependency set documented above. Provisioning writes the exact resolved distribution set to .dsflower_versions.txt; capabilities report the Flower/Torch/Opacus versions and the file's SHA-256. That post-install manifest is audit evidence, not a reproducible lock. A production administrator can set DSFLOWER_PYTHON_LOCK (or dsflower.python_lock) to a complete PyTorch requirements file with hashes for every transitive artifact, and DSFLOWER_NATIVE_TREE_PYTHON_LOCK (or dsflower.native_tree_python_lock) to a separate complete native-tree lock; provisioning then uses uv pip install --require-hashes and binds the lock SHA-256 into the venv marker. Never reuse either lock for the other dependency graph. Keep the same root-owned locks available for later health checks and re-provisioning. Set DSFLOWER_REQUIRE_PYTHON_LOCK=true and/or DSFLOWER_NATIVE_TREE_REQUIRE_PYTHON_LOCK=true to make an absent or invalid corresponding lock a fail-closed provisioning error. Set DSFLOWER_PYTHON_VERSION to an exact major.minor.patch to prevent interpreter patch drift; the flexible default 3.11 intentionally tracks a compatible patch release. Immutable container digests remain the deployment identity for byte-for-byte artifacts.

An existing OS-managed uv is part of the administrator's trusted computing base. If no uv is installed, dsFlower does not execute a mutable remote installer or a latest URL. Automatic Python bootstrap requires both an exact official release tag in DSFLOWER_UV_VERSION and its platform archive digest in DSFLOWER_UV_SHA256; a mismatch fails before extraction. For containers, persist /var/lib/dsflower/privacy/noise_root, its neighbourhood store and UUID pin/locks, the gated-release cache and the app-store directory. A secret-manager file may instead be selected through DSFLOWER_NODE_SECRET_FILE. Do not mount all of /var/lib/dsflower, because that path also contains the baked venvs.

The deterministic CSPRNG is keyed by the dedicated node secret and a canonical semantic release identity. This prevents averaging exact retries and stream reuse, but is a computational guarantee: disclosure of the persistent key can recreate streams for known semantic identities. Production nodes should keep it in a KMS/HSM-backed secret lifecycle. Preserving the same root preserves deterministic recomputation; rotating it intentionally starts a new randomness domain.

Binary association runner

The association track stages exactly one outcome and one exposure as a row-preserving 3x3 table (reference, positive, unknown). It performs one sticky joint-Gaussian release per node and allows only an all-or-nothing pooled result. Public typed levels and the row/patient estimand are bound by the association contract SHA-256; a second job SHA-256 binds that request to runner ABI 3, the exact runner hash and the public node roster. Neither hash contains a data path, run token, cohort size or privacy policy value.

The track uses the dependency-light Flower/NumPy runtime and has dedicated app entrypoints; it never falls back to the neural, HookApp, native-tree or validation runners. Association privacy retains the existing per-job/node accountant, with no historical privacy budget or rate limiter. Its complete released vector now uses the permanent neighbourhood store.

Server-side lifecycle

Stage Exported DataSHIELD methods
Connectivity/capabilities flowerPingDS, flowerCheckConnectivityDS, flowerGetCapabilitiesDS
Handle lifecycle flowerInitDS, flowerDestroyDS
Staging and validation flowerPrepareRunDS
SuperNode lifecycle flowerEnsureSuperNodeDS, flowerCleanupRunDS, flowerStatusDS
App integrity flowerAppPushDS, flowerAppInstallDS, flowerAppDeleteDS, flowerTier2PinDS
Privacy policy flowerPrivacyPolicyDS

See ARCHITECTURE.md for the precise trust boundary, mechanism contracts, deployment requirements and residual limitations.

Minimal researcher-side example

library(dsFlowerClient)
library(DSI)
library(DSOpal)

builder <- DSI::newDSLoginBuilder()
builder$append(server = "site1", url = "https://opal1.example.org",
               user = "researcher", password = "...",
               table = "PROJECT.training_data", driver = "OpalDriver")
conns <- DSI::datashield.login(builder$build(), assign = TRUE, symbol = "D")

# Bounds are public/domain-knowledge constants, not estimates queried from a node.
fit <- ds.flower.fit(
  conns, symbol = "D", target = "diagnosis",
  features = c("radius", "texture"), model = "pytorch_logreg",
  feature_bounds = list(lower = c(0, 0), upper = c(100, 1)),
  target_levels = c("control", "case")
)

DSI::datashield.logout(conns)

Authors

Barcelona Institute for Global Health (ISGlobal)

FedProx and local verification

The analyst may select dsFlowerClient::ds.flower.strategy.fedprox(mu = 0.01) for neural training or an admitted Hook. The node validates finite mu in [0,1] and, for positive neural mu, learning_rate * mu <= 1 throughout the public schedule. Neural contraction is applied after each DP optimizer step and L1 prox, using the fixed incoming round model. This is post-processing of DP/public values; clipping, gradients and the step accountant remain unchanged. Hook contraction is one update-level relaxation after the output gate, with eta=1; it does not instrument arbitrary Hook optimizers. A cache hit returns the final contracted bytes. Zero is exactly FedAvg on supported tracks. All five tree engines, association and standalone validation reject FedProx, including zero.

GitHub Actions workflows are removed in 0.7.1. Local R/Python suites, native verification scripts and integration harnesses remain available. A real federation is verified by a local multi-node integration harness. Use a private R library, private working/state directories and one harness session at a time; see the client local integration instructions.

Running the tests locally

The suites run from a source checkout (no hosted CI since 0.7.1). R, with the package and its dependencies installed:

R CMD INSTALL .
Rscript -e 'testthat::test_dir("tests/testthat", package = "dsFlower", load_package = "installed")'

Python suites use the node runtimes ($DSFLOWER_VENV_ROOT/pytorch and, for the native tree/XGBoost files, $DSFLOWER_VENV_ROOT/native-tree). Run each file under inst/python/tests/ with "$DSFLOWER_VENV_ROOT/pytorch/bin/python" -m pytest inst/python/tests/<file>.py; the few stand-alone script files are run directly with python inst/python/tests/<file>.py and report through their exit status. Tests that need a node secret use a throwaway one when DSFLOWER_TEST_ALLOW_EPHEMERAL_SECRET=1; tests that need optional components (dsImaging, the curated XGBoost bundle) skip when those are absent. From a dsFlowerClient checkout next to this one, python3 tools/check-runner-sync.py --server ../dsFlower must print that both runner copies are byte-identical. The retained XGBoost release verifier is native/xgboost/tests/real_runner_e2e.py (native-tree runtime, bundle configured).

About

To be supplied

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages