Repository navigation
Publish consolidated BackIntel demo proof of concept - #2
hudsonaikins wants to merge 52 commits into
Conversation
|
@codex review |
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b6fbe0895c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| means = metric_values(cited, {'mean'}) | ||
| counts = metric_values(cited, {'count', 'labeled', 'records', 'sample_size', 'source_size'}) |
There was a problem hiding this comment.
Separate prior metrics from current claim validation
When compare_snapshots returns different current and previous values, metric_values adds both snapshots' means and counts to the same allowance. For example, with a current mean of 0.2 and previous mean of 0.8, a factual claim that the current observed rate is 80% passes validation and can be published as the goal's latest successful answer. Preserve the snapshot role when validating narrative numbers so prior values cannot substantiate current-snapshot claims.
Useful? React with 👍 / 👎.
| unresolved = c.execute('SELECT count(*) FROM backintel.analysis_requests WHERE run_id=%s AND charge IS NULL', (run_id,)).fetchone()[0] | ||
| if unresolved: | ||
| raise RuntimeError('Reconcile uncertain charges before resuming this run') | ||
| resumed=c.execute("UPDATE backintel.capability_jobs SET state='queued',cancel_requested=false,result_sha256=NULL,due_at=now(),error=NULL,lease_until=NULL,max_attempts=GREATEST(max_attempts,attempts+3) WHERE job_id=%s AND state IN ('completed','cancelled','failed') RETURNING job_id", (job_id,)).fetchone() |
There was a problem hiding this comment.
Cap resumed jobs at the schema attempt limit
After repeated partial runs, this expression can exceed the database constraint: an initial partial attempt followed by two resumed partial attempts leaves attempts = 3, so the next resubmission tries to set max_attempts = 6, while the schema permits at most 5. PostgreSQL then rejects the submission and the deterministic run remains permanently partial. The fresh recovery update now includes failed jobs, but it should cap or explicitly reject retries at the five-attempt boundary rather than violating the constraint.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| key = digest({"task":task_record["sha256"],"training":[[r["feature"]["sha256"],r["outcome"]["sha256"]] for r in training], | ||
| "route":route,"feature_set":feature_set,"at":at,"implementation_mode":"real","libraries":libraries,"config":config}) |
There was a problem hiding this comment.
Bind real-model cache keys to implementation bytes
When predictor code changes while the persisted model directory, training evidence, configuration, and library versions remain unchanged, this key is unchanged, so prepare_real returns the prior evidence record or reloads the old package instead of executing the new implementation. A deployment can therefore evaluate and publish results from artifacts built by different code than the reviewed candidate; include an implementation fingerprint in the cache identity and manifest, as the newer analysis comparison path already does.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| tasks = store.list("task") | ||
| if not tasks: | ||
| raise ValueError("Task has no registered contract") | ||
| task = max(tasks, key=lambda r: r["available_at"]) |
There was a problem hiding this comment.
Select task revisions unambiguously when issuing grants
When a revised task contract is registered with the default available_at=0, both versions tie here; because evidence is ordered by identity, max(..., key=available_at) selects the first tied hash rather than the most recently registered contract. The resulting token can be bound to the old task SHA while get_artifact returns the new version's report, causing scoped_artifact to deny access, and the packaging path repeats the same stale-version selection. Require monotonic revision times or persist an explicit active task version.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b1ff3c1d74
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
|
||
| def predict(model, rows): | ||
| import joblib | ||
| body=model['body']; directory=model_root()/'Analysis'/body['id'] |
There was a problem hiding this comment.
Reject active models prepared by a different runtime
When analysis_models.py, its preprocessing logic, or a recorded library version changes after promotion, this path validates only the serialized artifact hash and continues scoring with the current matrix_features/decide code. Although comparison manifests record dependencies.implementation_sha256 and library versions, predict never compares them with the running environment, so a deployed update can publish predictions from a code/artifact combination that was never evaluated or approved; verify the recorded runtime dependencies before loading the active model.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| g = db.goal(identity) | ||
| with db.connect() as c, c.transaction(): | ||
| run_id = _admit_run(c, g, actor, question, operation) |
There was a problem hiding this comment.
Lock the current goal row before deriving a run
If a manager revises or pauses a goal after this read but before _admit_run inserts the run, the request is accepted using the stale version, question, confirmation state, and active model. The worker then reaches analysis_store.check_run and cancels that newly accepted work because its goal_version no longer matches, instead of either admitting the current goal or rejecting the concurrent edit at submission time; reload and lock the goal inside this transaction before deriving the run identity.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| existing = store.find("real_plan", "history-v1") | ||
| if existing: |
There was a problem hiding this comment.
Bind cached real plans to the generated scenario data
When a persisted demo task is reused after history() or its scenario configuration changes, the normal real_prepare path passes no history_data, so this returns the old history-v1 plan without comparing any dataset identity; prepare_followups has the same gap when event_data is omitted. A new candidate can therefore run paid interpretation, model comparison, and packaging against rows and follow-up events prepared by an earlier implementation while being presented as the current run. Always derive and hash the default history/events before accepting these cached plans.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| useEffect(()=>{ | ||
| if (!workspace?.demo) return | ||
| const interval=setInterval(()=>setRefresh(value=>value+1),1500) | ||
| return ()=>clearInterval(interval) | ||
| },[Boolean(workspace?.demo)]) |
There was a problem hiding this comment.
Stop polling after the demo reaches a terminal state
Whenever a workspace contains any demo object, this installs a 1.5-second refresh loop regardless of whether its status is running, completed, blocked, or unavailable. The shipped recorded package already contains a completed demo, so every idle browser tab continuously rereads the packet and SQLite review state and rerenders the entire workspace indefinitely; restrict polling to statuses that can still advance or stop the interval on terminal states.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| sent = db.query("UPDATE backintel.analysis_requests SET status='sent' WHERE id=%s AND status='reserved' RETURNING id", (call_id,), one=True) | ||
| if not sent: | ||
| raise RuntimeError('Provider reservation changed before dispatch; no request was sent') | ||
| try: | ||
| with httpx.Client(timeout=120) as client: | ||
| response = client.post('https://openrouter.ai/api/v1/responses', headers={'Authorization': 'Bearer '+key}, json=payload) |
There was a problem hiding this comment.
Recheck cancellation before dispatching the paid request
If a manager cancels a running analysis after reserve verifies the job but before this transition, jobs.cancel sets cancel_requested while leaving the running job state unchanged, and this update still marks the request sent and immediately performs the provider POST. The run is later cancelled at acceptance, but the supposedly cancelled operation has already incurred a charge; make the reserved-to-sent transition conditional on the associated job still being uncancelled, or recheck check_run atomically before dispatch.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 512c37916c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| raise ValueError('Requested predictor was not validated in this comparison') | ||
| if locked[1] and locked[0].get('approved_route') != route: | ||
| raise ValueError('A promoted candidate has an immutable predictor route') | ||
| run_id = _admit_run(c, {**g, 'active_model': identity}, actor) |
There was a problem hiding this comment.
Admit the refresh with the promoted model
Every promotion queues its refresh against the previous active model because _admit_run immediately reloads the goal row and discards the active_model: identity override passed here, while the goal is updated only afterward. On the first promotion the run therefore stores model_id: null; on later promotions it stores the prior model, so predict cannot evaluate the approved candidate and the completed refresh remains stale relative to the goal's new active_model. Activate the candidate before admission or pass an explicit model override that survives the locked reload.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| prefix = "-real-" if operation.startswith("real_") else "-" | ||
| job_id = enqueue(Evidence(connection,f"{config['id']}{prefix}{demo_id}"),payload,request_id) | ||
| try: | ||
| return {"job_id":job_id,"result":await asyncio.to_thread(execute,job_id)} |
There was a problem hiding this comment.
Preserve task ordering for directly submitted jobs
When two graph requests for the same task overlap, this executes each newly enqueued job directly, bypassing the earlier-job predicate in jobs.runnable. Both jobs can therefore be claimed as running, and whichever worker acquires the task advisory lock first executes first; a later dependent stage such as real_compare can fail before an earlier interpretation finishes, despite the queue explicitly promising that no task bypasses an earlier queued, retrying, or failed job. Enforce the sequence constraint in claim or route these submissions through the ordered dispatcher.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 48a4f6088b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if domain=='maintenance': | ||
| body['unit_mapping']='dataset-qualified-engine-v2' | ||
| identity=digest(body) | ||
| db.save_snapshot(domain, identity, body, rows) |
There was a problem hiding this comment.
Lock the source before adapting a refresh
When a changed-file import is adapting while a manager correction commits, _import_source performs its read and adaptation before acquiring the source advisory lock; this save_snapshot then waits for the correction and replaces latest_snapshot with rows based on the pre-correction state. The correction silently becomes non-current and scheduled analyses use the original value, so acquire the source lock before deriving the import or recheck/rebase the latest snapshot under the lock before publishing.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| if g['paused'] or g['version'] != r['goal_version']: | ||
| raise PermissionError('Goal is paused or this run was superseded') |
There was a problem hiding this comment.
Stop runs after goal confirmation is revoked
When a manager sends confirmed: false without changing the goal body, revise_goal leaves the version unchanged, and this eligibility check only considers pause and version state. An already-running analysis therefore continues paid calls and can publish itself as last_success even though its definitions are no longer confirmed; include the current confirmation state in check_run.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| result = import_source(r['body']['domain'], {'id': r['owner']}) | ||
| db.write('UPDATE backintel.analysis_runs SET snapshot_id=%s WHERE id=%s',(result['snapshot'],identity)) | ||
| result['goals'] = schedule_snapshot(r['body']['domain'], result) |
There was a problem hiding this comment.
Recheck cancellation before publishing an import
When a manager cancels a running import after the initial check_run, import_source commits the new snapshot through separate database connections and schedule_snapshot queues dependent analyses before the next cancellation check. The job is ultimately reported as cancelled, but the source has already changed and the dispatcher can run the newly queued analyses, including paid work; revalidate the run under the source lock before publishing and scheduling the import.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| existing = store.connection.execute("SELECT kind,payload,repeat_seconds FROM backintel.capability_triggers WHERE trigger_id=%s",(trigger_id,)).fetchone() | ||
| if existing: | ||
| if existing != (kind,payload,repeat_seconds): | ||
| raise ValueError("Conflicting trigger identity") |
There was a problem hiding this comment.
Include occurrence count in trigger replay checks
When the same task and request ID are scheduled again with the same kind, payload, and interval but a different occurrences value, this lookup treats the request as an idempotent replay and silently returns the old trigger. The caller can therefore request five executions but retain a previously created one-execution trigger, or vice versa; persist or derive the original occurrence count and include it in the conflicting-identity check.
Useful? React with 👍 / 👎.
| if not row or set(body.features)-set(row['features']): | ||
| raise ValueError('Unknown record or feature') | ||
| row['features'].update(body.features) | ||
| row['groups'].update(feature_groups(domain, row['features'])) |
There was a problem hiding this comment.
Preserve feature types in source corrections
When a manager corrects a numeric feature such as age or tenure with a JSON string, or a categorical feature with a number, CorrectionInput accepts the value and this update persists it without checking the original/domain schema. matrix_features branches on the Python value type, so an approved predictor silently interprets the corrected field as a different categorical or numeric feature rather than rejecting the malformed correction; validate each replacement against the field's established type before saving the snapshot.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4ba5f1fb21
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if f['kind'] == 'fact' and any(r.get('kind') == 'estimate' for r in results if r['evidence_id'] in f['evidence_ids']): | ||
| raise ValueError('Predicted evidence cannot support a factual finding') |
There was a problem hiding this comment.
Require estimates to cite prediction evidence
When a run contains both summarize and predict results, an estimate finding can cite only the observed summary because this check is asymmetric: it rejects facts backed by estimates but accepts estimates backed by observed evidence. For example, a 20% observed mean and an 80% prediction allow “risk is 20%” to pass as an estimate, while the later global prediction check is satisfied merely because predict was called somewhere in the run. Require estimate findings to cite evidence whose kind is estimate so observed outcomes cannot be published as forecasts.
Useful? React with 👍 / 👎.
| except (OSError, ValueError, PermissionError, RuntimeError) as error: | ||
| db.write('UPDATE backintel.analysis_sources SET body=body || %s WHERE id=%s', | ||
| (Jsonb({'last_refresh_error':safe_error(error),'last_checked_at':time.time()}),domain), connection=connection) |
There was a problem hiding this comment.
Persist refresh failures outside the rolled-back import transaction
When an API-queued import fails, handle passes store.connection and wraps the call in a transaction; this error update therefore runs on that same transaction and is rolled back when the exception is re-raised. The run becomes partial, but last_refresh_error and last_checked_at are never retained, so /api/v1/goals continues reporting the previous freshness instead of refresh_failed. Record the failure after the import transaction has rolled back, or use a separate committed connection.
Useful? React with 👍 / 👎.
| old = {r['id'] for r in db.records(latest['snapshot_id']) if r['target'] is not None and r['split']=='train'} | ||
| new = {r['id'] for r in db.records(update['snapshot'], connection=connection) if r['target'] is not None and r['split']=='train'} | ||
| attempt = db.query("SELECT created_at FROM backintel.analysis_runs WHERE goal_id=%s AND body->>'operation'='training' ORDER BY created_at DESC LIMIT 1", (g['id'],), one=True) | ||
| interval = time.time()-(attempt or latest)['created_at'].timestamp() | ||
| if len(new-old)>=CONFIG['limits']['new_labels'] and interval>=CONFIG['limits']['candidate_interval_seconds']: |
There was a problem hiding this comment.
Detect modified training rows during snapshot refresh
When refreshed source bytes alter or remove existing training rows without adding at least 128 new IDs, these sets remain identical (or new-old stays below the threshold), so no replacement candidate is trained and the active model is not invalidated. Subsequent analyses can therefore be marked current while scoring the new snapshot with a model trained and evaluated against the old feature/target values; the explicit correction path already invalidates models for this case. Compare the relevant row contents as well as IDs, and invalidate or retrain models affected by modified or removed preparation rows.
Useful? React with 👍 / 👎.
| row['features'].update(body.features) | ||
| row['groups'].update(feature_groups(domain, row['features'])) | ||
| if body.target is not None: | ||
| if CONFIG['sources'][domain]['kind']=='classification' and body.target not in (0,1): | ||
| raise ValueError('Classification target must be 0 or 1') | ||
| row['target']=body.target |
There was a problem hiding this comment.
Reject corrections that do not change a source value
A correction containing only the required record_id and explanation fields defaults to features={} and target=None, so neither branch changes the row, yet the endpoint still creates a new snapshot, invalidates every model whose split contains the record, clears an active model, and schedules refreshes. An explicit target: null has the same effect. Require at least one actual value change, or distinguish an explicitly requested target removal from an omitted target, before publishing and invalidating models.
Useful? React with 👍 / 👎.
| @lru_cache(maxsize=2) | ||
| def _load(root_value: str, model_json: str): |
There was a problem hiding this comment.
Recheck model-use approval on every prediction
After the first prediction for a model, _load returns its cached estimator without executing _allow_model_use again. If an operator deletes or revokes model-use-approval.json while the service remains alive, subsequent predictions with the same model continue running and can feed refreshed reports despite the explicit local execution approval no longer existing. Keep estimator loading cached, but move the approval check outside the cached function or bind cache validity to a current approval identity.
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c8813fd078
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| for label in labels: | ||
| if any(character.isalpha() for character in label): | ||
| text=re.sub(r'(?<!\w)'+re.escape(label)+r'(?!\w)','entity',text,flags=re.I) | ||
| for match in re.finditer(r'(?<!\w)-?\d[\d,]*(?:\.\d+)?', text): |
There was a problem hiding this comment.
Validate qualitative comparisons against cited rows
When a cited summary contains groups A at 0.2 and B at 0.8, a factual finding such as “A has the highest rate” passes because validation examines only numeric literals. The claim is then copied into the primary summary on line 105, so users can receive a conclusion that directly contradicts the cited table. Validate group rankings and comparative language, or construct these claims from the typed rows.
Useful? React with 👍 / 👎.
| snapshot = 'previous' if snapshots and snapshots[-1].lower() in ('previous', 'prior', 'historical') else 'current' | ||
| metrics = [r.get(snapshot) if r.get('tool') == 'compare_snapshots' or 'current' in r and 'previous' in r else r for r in cited] | ||
| means = metric_values(metrics, {'mean'}) | ||
| counts = metric_values(metrics, {'count', 'labeled', 'records', 'sample_size', 'source_size'}) |
There was a problem hiding this comment.
Allow cited missingness counts during answer validation
For an inspect_source result such as missing: {"age": 12}, a supported claim like “12 records are missing age” is rejected: recursive metric extraction ignores the value because it is keyed by the feature name rather than one of these count keys. This makes source-quality questions fail after a paid call unless the missing count happens to equal another allowed count; recognize values under the typed missing mapping as missingness counts.
Useful? React with 👍 / 👎.
| source = db.source(domain, connection=c) | ||
| if not source['body'].get('terms_acknowledged'): | ||
| raise PermissionError('Source access and terms must be confirmed before import') |
There was a problem hiding this comment.
Recheck source terms before publishing an import
When a manager revokes source terms while a changed-file import is adapting, this one initial check has already passed, and the later check_run verifies only the owner grant and cancellation state. Because the terms endpoint does not acquire the source advisory lock, the import can still publish and schedule analyses while the source record says access is no longer acknowledged; synchronize terms updates with this lock and revalidate the current terms state before saving the snapshot.
Useful? React with 👍 / 👎.
| @lru_cache(maxsize=2) | ||
| def _load(root_value: str, model_json: str): |
There was a problem hiding this comment.
Recheck model-use approval on cached predictions
After a model has been loaded once, changing or removing model-use-approval.json has no effect for the same model because this cache returns the estimator without rerunning _allow_model_use. A long-lived worker can therefore continue executing actual predictions after the operator revokes the explicit model-use approval; keep deserialization cached if needed, but check the current approval before every prediction.
Useful? React with 👍 / 👎.
| row=next((r for r in rows if r['id']==body.record_id),None) | ||
| if not row or set(body.features)-set(row['features']): | ||
| raise ValueError('Unknown record or feature') |
There was a problem hiding this comment.
Reject corrections that do not change the record
A valid CorrectionInput may omit both features and target, and this check accepts it as long as the record exists. The endpoint then creates a new snapshot and invalidates any model whose split contains that record even though no source value changed; identical replacement values have the same effect. Require at least one actual field or target change, or return an unchanged result before snapshot creation and model invalidation.
Useful? React with 👍 / 👎.
| for label in labels: | ||
| if any(character.isalpha() for character in label): | ||
| text=re.sub(r'(?<!\w)'+re.escape(label)+r'(?!\w)','entity',text,flags=re.I) | ||
| for match in re.finditer(r'(?<!\w)-?\d[\d,]*(?:\.\d+)?', text): |
There was a problem hiding this comment.
Parse complete scientific-notation values
When the analyst writes a supported value in scientific notation, this pattern validates only its mantissa. For example, 1e-3 is treated as 1, so a legitimate probability of 0.001 is rejected, or the claim can pass incorrectly if an unrelated cited count equals 1. Match the exponent and require a token boundary before comparing the parsed value with typed metrics.
Useful? React with 👍 / 👎.
| for item in manifest["files"]: | ||
| path = (root / item["path"]).resolve() | ||
| if not path.is_relative_to(root.resolve()) or not path.is_file() or sha(path) != item["sha256"]: | ||
| raise ValueError(f"Package file is missing or changed: {item['path']}") |
There was a problem hiding this comment.
Reject files omitted from the package manifest
When an extracted or redistributed package contains an additional file that is absent from Manifest.json, --check still reports success because verification only walks declared entries. This contradicts the package's stated whole-package verification and leaves additions under served or source directories undetected; compare the actual immutable package file set with the manifest, excluding only the documented mutable .demo-state area and the manifest itself.
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e3a824e400
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if question is not None and (not isinstance(question, str) or not 1 <= len(question.strip()) <= 2000): | ||
| raise ValueError('Invalid follow-up question') | ||
| question = question or g['body']['question'] | ||
| run_id = digest([identity, g['version'], snapshot, operation, question, g['active_model']]) |
There was a problem hiding this comment.
Include the analysis implementation in run identities
When the analyst prompt, validation/tool logic, or config/analysis.json model settings change while the goal, snapshot, question, and active model remain unchanged, this digest still selects the old run. The following conflict-tolerant insert and enqueue therefore reuse its completed job, so running the analysis after deployment returns the pre-deployment result and never exercises the fix until unrelated data or goal state changes. Add an implementation/config fingerprint to the run identity.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| if standing: | ||
| connection.execute('UPDATE backintel.analysis_goals SET last_success=%s WHERE id=%s AND version=%s', (identity, g['id'], r['goal_version'])) |
There was a problem hiding this comment.
Recheck goal eligibility while publishing success
When a manager pauses the goal or revokes confirmation after the final check_run but before this transaction, neither edit increments the goal version, so this update still installs the result as last_success. Although check_run now examines confirmation, the success transaction does not lock the goal and this predicate checks only the version; make publication conditional on the goal still being confirmed and unpaused under the same transaction.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
|
||
| def catalog(): | ||
| for domain, spec in CONFIG['sources'].items(): | ||
| write('INSERT INTO backintel.analysis_sources(id,domain,body) VALUES(%s,%s,%s) ON CONFLICT DO NOTHING', |
There was a problem hiding this comment.
Invalidate source consent when registered terms change
When a deployment changes a source's configured files, license, Kaggle identity, or competition status, an existing database row is left untouched by this conflict clause. Imports nevertheless use the new CONFIG entry while authorization checks only the row's old terms_acknowledged boolean, so consent granted for the previous source specification silently authorizes the replacement and the UI continues displaying stale terms. Persist a source-spec/terms fingerprint and require acknowledgment again whenever it changes.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| c.execute('INSERT INTO backintel.analysis_models(id,goal_id,snapshot_id,body) VALUES(%s,%s,%s,%s) ON CONFLICT DO NOTHING', | ||
| (candidate, g['id'], r['snapshot_id'], Jsonb(manifest))) |
There was a problem hiding this comment.
Publish model candidates only with the accepted training run
When cancellation or another failure occurs after the eligibility check but before model_job returns, this separate connection has already committed the candidate even though the enclosing job is later cancelled or partial. The models endpoint still lists that orphan as awaiting approval, and promote does not require a succeeded originating run, so a manager can activate a candidate whose training result was never accepted. Insert the candidate through the job transaction or persist and verify its successful run provenance before promotion.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2f1302702b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| db.authorize(actor, g['domain'], ('manager',) if operation == 'training' else ('manager', 'analyst')) | ||
| if not g['confirmed'] or g['paused']: | ||
| raise ValueError('Confirm the goal definitions and enable the goal first') | ||
| snapshot = db.source(g['domain'], connection=c)['latest_snapshot'] |
There was a problem hiding this comment.
Block analysis after source terms are revoked
When a manager sets terms_acknowledged to false after a snapshot exists, this admission path uses only latest_snapshot, so analysts can still start analysis or training against the retained data; analysis_store.check_run likewise never rechecks the source terms, allowing an already-running job and paid provider calls to continue. Require current acknowledged terms both when admitting work and at execution/publication checkpoints.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| locked = c.execute('SELECT body,promoted FROM backintel.analysis_models WHERE id=%s FOR UPDATE', (identity,)).fetchone() | ||
| if locked[0].get('invalidated_by'): | ||
| raise ValueError('This candidate was invalidated by a source correction') |
There was a problem hiding this comment.
Reject candidates from superseded snapshots
If candidate comparison completes on snapshot A and a normal source refresh publishes snapshot B before manager approval, the candidate has no invalidated_by marker and passes this check. Promotion then activates the model and immediately analyzes snapshot B even though the manager-reviewed evaluation belongs to A; compare the candidate's snapshot_id with the source's current snapshot under the source lock before activation.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| for g in result: | ||
| source=db.source(g['domain']) | ||
| previous=db.run(g['last_success']) if g['last_success'] else None | ||
| g['freshness']='refresh_failed' if source['body'].get('last_refresh_error') else 'current' if previous and previous['snapshot_id']==source['latest_snapshot'] and previous['goal_version']==g['version'] and previous['body'].get('model_id')==g['active_model'] else 'stale' if previous else 'unanswered' |
There was a problem hiding this comment.
Mark answers stale after analysis implementation changes
After an analyst prompt, validator, tool, or config/analysis.json change, this still labels the previously published result current whenever its snapshot, goal version, and model match. Although _admit_run now includes analysis_identity() in new run IDs, that fresh evidence is not recorded or checked here, so the workspace can present a pre-deployment answer as current until another run happens; persist the run's implementation identity and include it in freshness evaluation.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| source=db.source(domain) | ||
| rows=db.records(source['latest_snapshot']) |
There was a problem hiding this comment.
Reuse the locked connection when loading correction sources
When a deployment changes a source specification before the catalog row has been refreshed, this call opens a second connection, detects the specification mismatch, and invokes catalog, which waits for the same source advisory lock already held by c. The correction request therefore hangs indefinitely instead of revoking the old consent and returning an error; pass connection=c through source and the dependent reads so catalog refresh is reentrant within this transaction.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| {view==='goals'&&<section><h2>Save a question</h2>{manager?<form onSubmit={e=>{e.preventDefault();action(async()=>{const g=await api<Goal>('/goals','POST',{domain,question});setGoalId(g.id);});}}><label htmlFor="question">Question to monitor</label><textarea id="question" value={question} onChange={e=>setQuestion(e.target.value)} maxLength={2000} required/><button disabled={busy}>Propose goal</button></form>:<p>Managers configure standing questions. Select a permitted goal to investigate.</p>}{selected&&<div className="definitions"><h3>Confirm business meaning</h3><dl>{Object.entries(selected.body.definitions).map(([k,v])=><React.Fragment key={k}><dt>{k}</dt><dd>{v}</dd></React.Fragment>)}</dl>{manager&&<><form key={selected.id+':'+selected.version} onSubmit={e=>{e.preventDefault();const data=new FormData(e.currentTarget);action(async()=>{await api('/goals/'+goalId,'PATCH',{threshold:Number(data.get('threshold')),budget_usd:Number(data.get('budget'))});});}}><label htmlFor="threshold">Notify when a group mean changes by at least</label><input id="threshold" name="threshold" type="number" min="0" max="100000" step="any" defaultValue={selected.body.notification_delta} required/><label htmlFor="budget">Maximum frontier cost per run (USD)</label><input id="budget" name="budget" type="number" min="0.01" max="1" step="0.01" defaultValue={selected.body.budget_usd||1} required/><button disabled={busy}>Save goal limits</button></form><label htmlFor="edit-question">Standing question</label><textarea id="edit-question" value={editQuestion} onChange={e=>setEditQuestion(e.target.value)}/><div className="actions"><button disabled={busy||editQuestion===selected.body.question} onClick={()=>action(async()=>{await api('/goals/'+goalId,'PATCH',{question:editQuestion});})}>Save revised question</button><button disabled={busy||selected.confirmed} onClick={()=>action(async()=>{await api('/goals/'+goalId,'PATCH',{confirmed:true});})}>Confirm definitions</button><button disabled={busy} onClick={()=>action(async()=>{await api('/goals/'+goalId,'PATCH',{paused:!selected.paused});})}>{selected.paused?'Resume goal':'Pause goal'}</button></div></>}</div>}</section>} | ||
| {view==='analysis'&&<section><div className="section-heading"><h2>{selected?.body.question||'Choose a saved goal'}</h2>{selected&&canAnalyze&&<button disabled={busy||!selected.confirmed||selected.paused} onClick={()=>action(async()=>{await api('/goals/'+goalId+'/runs','POST',{});})}>Run analysis</button>}</div>{pending&&<p role="status" className="progress">{pending.body.operation==='training'?'Comparing prediction methods':'Analysis in progress'} · {progress||pending.status}</p>}{selected&&<p className="muted">Answer status: {selected.freshness} · Goal version {selected.version}</p>}{current?.result?<><p className="answer">{current.result.summary}</p>{current.result.findings?.map((f,i)=><article className="finding" key={i}><span className="tag">{f.kind}</span><p>{f.claim}</p>{f.evidence_ids.map(id=><button className="text-button" key={id} onClick={()=>action(async()=>{setEvidence(await api('/evidence/'+id));})}>Inspect calculation</button>)}</article>)}{current.result.tables?.map((t,i)=><div key={i}><h3>{t.title}</h3><div className="table-scroll"><table><thead><tr><th>Group</th><th>Records</th><th>Known values</th><th>Mean</th></tr></thead><tbody>{t.rows.map(row=><tr key={row.group}><td>{row.group}</td><td>{format(row.count)}</td><td>{format(row.labeled)}</td><td>{format(row.mean)}</td></tr>)}</tbody></table></div><figure aria-label={t.title+' mean comparison'}>{t.rows.filter(r=>r.mean!==null).map(row=><div className="chart-row" key={row.group}><span>{row.group}</span><meter min={0} max={Math.max(1,...t.rows.map(r=>r.mean||0))} value={Math.max(0,row.mean||0)} aria-label={row.group+' mean'}/><span>{format(row.mean)}</span></div>)}<figcaption>Calculated values for the displayed groups.</figcaption></figure></div>)}{current.result.limitations?.map((s,i)=><p className="caveat" key={i}>{s}</p>)}<p className="muted">Measured inference: ${format(current.result.usage.provider_usd)} · {current.result.usage.charge_status}</p></>:<div className="empty"><h3>No completed answer yet</h3><p>Select a goal, confirm its definitions, and run the analysis. Partial runs remain in history.</p></div>}</section>} | ||
| {view==='conversation'&&<section><h2>Ask a follow-up</h2><p>The standing goal changes only when a manager saves a revision.</p><form onSubmit={e=>{e.preventDefault();action(async()=>{await api('/goals/'+goalId+'/runs','POST',{question:followup});setFollowup('');});}}><label htmlFor="followup">Follow-up question</label><textarea id="followup" value={followup} onChange={e=>setFollowup(e.target.value)} required maxLength={2000}/><button disabled={busy||!canAnalyze||!selected?.confirmed}>Investigate</button></form>{runs.filter(r=>r.body.operation==='analysis'&&r.body.question!==selected?.body.question).map(r=><article className="finding" key={r.id}><h3>{r.body.question}</h3><p className="muted">{r.status}</p>{r.result&&<><p className="answer">{r.result.summary}</p>{r.result.findings?.map((f,i)=><div key={i}><p>{f.claim}</p>{f.evidence_ids.map(id=><button className="text-button" key={id} onClick={()=>action(async()=>{setEvidence(await api('/evidence/'+id));})}>Inspect calculation</button>)}</div>)}{r.result.limitations?.map((text,i)=><p className="muted" key={i}>{text}</p>)}</>}</article>)}</section>} | ||
| {view==='review'&&<section><h2>Prediction candidates</h2><p>Training can run automatically. A manager approves each replacement.</p>{manager&&selected&&<button disabled={busy||!selected.confirmed||selected.paused} onClick={()=>action(async()=>{await api('/goals/'+goalId+'/runs','POST',{operation:'training'});})}>Compare real predictors</button>}{models.map(m=><article className="finding" key={m.id}><p className="muted">Candidate {m.id.slice(0,12)} · {m.promoted?'Approved':'Awaiting manager'}</p><div className="table-scroll"><table><thead><tr><th>Method</th><th>Features</th><th>Measured result</th></tr></thead><tbody>{m.body.methods.map((method,i)=><tr key={i}><td>{method.route}</td><td>{method.features||'Baseline'}</td><td>{Object.entries(method.metrics).filter(([,v])=>typeof v==='number').map(([k,v])=>k+': '+format(v as number)).join(' · ')}</td></tr>)}</tbody></table></div>{manager&&!m.promoted&&m.body.methods.filter(method=>['catboost','tabiclv2'].includes(method.route)).map(method=><button key={method.route+'-'+method.features} disabled={busy} onClick={()=>action(async()=>{await api('/models/'+m.id+'/promote','POST',{route:method.route+'-'+(method.features||'facts')});})}>Approve {method.route} ({method.features||'facts'})</button>)}</article>)}<h2>Review notes and corrections</h2>{canAnalyze&&selected&&<form onSubmit={e=>{e.preventDefault();action(async()=>{await api('/goals/'+goalId+'/reviews','POST',{text:note,kind:'correction'});setNote('');});}}><label htmlFor="review-note">Explain the finding to review</label><textarea id="review-note" value={note} onChange={e=>setNote(e.target.value)} required/><button disabled={busy}>Save review note</button></form>}{reviews.map(v=><article className="finding" key={v.id}><span className="tag">{v.body.kind}</span><p>{v.body.text}</p><small>{v.actor}</small></article>)}{manager&&<details><summary>Apply a source target correction</summary><form onSubmit={e=>{e.preventDefault();action(async()=>{await api('/sources/'+domain+'/corrections','POST',{record_id:correctionId,target:Number(correctionTarget),explanation:note||'Manager correction'});});}}><label htmlFor="record-id">Record identifier</label><input id="record-id" value={correctionId} onChange={e=>setCorrectionId(e.target.value)} required/><label htmlFor="corrected-target">Corrected target value</label><input id="corrected-target" type="number" step="any" value={correctionTarget} onChange={e=>setCorrectionTarget(e.target.value)} required/><button disabled={busy}>Apply correction and invalidate affected findings</button></form></details>}</section>} |
There was a problem hiding this comment.
Select predictor routes without consuming the test metrics
On the manager review screen, each approval button is displayed beside method.metrics, which analysis_models.compare computes from the nominal test split, while the separate calibration_metrics are not shown. A manager choosing the route from this screen therefore uses the test outcomes for model selection, so the subsequently reported test performance is no longer held out; use calibration evidence or a predeclared selection rule for approval and reserve test metrics for evaluation after the route is fixed.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
| for name in ("Validation.json", "Replay.json", "Budget.json", "NumericRoundingChecks.json", "ResumeApproval.json", "AuthorizationClosure.Resumed.json", "ReactInterpretationRunning.png", "ReactReviewed.png", "ReflexReviewed.png"): | ||
| shutil.copyfile(run / name, records / name) | ||
| shutil.copyfile(receipt_path, records / "Integration.json") | ||
| shutil.copytree(receipt_path.parent / "Package", records / "ExecutionPackage") |
There was a problem hiding this comment.
Exclude audience bearer tokens from recorded packages
When the recorded business archive is built within the one-hour grant lifetime, this copies the entire capability package, including the explicitly private AudienceAccess.json, into the distributable ZIP. The original file's 0600 mode no longer protects its contents inside the tracked/readable archive, so anyone who can read the package can use correction-capable bearer tokens against a concurrently running local audience server; omit the private grant file when copying execution evidence.
AGENTS.md reference: AGENTS.md:L4-L4
Useful? React with 👍 / 👎.
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 631f31b339
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if actor['role']=='viewer': | ||
| return [{**{key:m[key] for key in ('id','goal_id','snapshot_id','promoted','created_at')}, | ||
| 'body':model_summary(m['body'])} for m in candidates] | ||
| return candidates |
There was a problem hiding this comment.
Withhold test metrics until route approval
Although the UI now renders calibration_metrics for unpromoted candidates, this endpoint still returns the complete candidate body—including held-out metrics, test predictions, and split membership—to the manager who chooses the route. A manager or API client can therefore inspect test outcomes before promotion and select the model using them, invalidating the subsequent held-out evaluation; return only calibration evidence until the route is fixed.
Useful? React with 👍 / 👎.
| if manifest.exists(): | ||
| prior=json.loads(manifest.read_text()) | ||
| for artifact in prior['artifacts']: | ||
| if file_sha(directory/artifact['file']) != artifact['sha256']: | ||
| raise ValueError('Model artifact identity mismatch') | ||
| return prior |
There was a problem hiding this comment.
Bind cached manifests to the requested comparison
When an existing comparison directory contains a modified or stale manifest.json, this branch validates only the artifact hashes named by that manifest and then trusts all of its other contents. Changing claimed metrics, splits, dependencies, snapshot identity, or even the manifest ID leaves the model files untouched and is accepted as manager-review evidence; verify the cached manifest against the freshly computed identity/domain/snapshot/dependencies and expected cohort before returning it.
Useful? React with 👍 / 👎.
| changed = c.execute('UPDATE backintel.analysis_goals SET active_model=%s WHERE id=%s AND version=%s AND confirmed AND NOT paused RETURNING id', (identity, g['id'], g['version'])).fetchone() | ||
| if not changed: | ||
| raise ValueError('Goal changed during promotion; reload it and try again') | ||
| run_id = _admit_run(c, g, actor) |
There was a problem hiding this comment.
Refresh the standing answer when reactivating a model
When candidate A was promoted previously, candidate B is active now, and a manager re-promotes A with the same route, this update succeeds but _admit_run derives A's original run ID and reuses its already-completed job. The dispatcher has nothing runnable, while last_success still points to B, so the goal remains permanently stale and even a normal rerun reuses the same completed A job; either restore A's exact prior result as last_success atomically or admit a distinct refresh run for reactivation.
Useful? React with 👍 / 👎.
| raise KeyError("Case not found") | ||
| item = deepcopy(self.cases[case_id]) | ||
| with self.connect() as connection: | ||
| rows = connection.execute("SELECT decision,reason,revision,recorded_at FROM reviews WHERE case_id=? AND source_sha256=? ORDER BY revision DESC", (case_id, item["evidence"]["sha256"])).fetchall() |
There was a problem hiding this comment.
Bind reviews to the complete decision context
Reviews are keyed only by the source-evidence hash, so the live business-demo watcher can replace a case's finding or prediction while retaining the same source lineage and this query will attach the old review to the new analysis. A decision recorded while interpretation and estimates were still pending can therefore appear as approval of later model output and unlock its later outcome without any stale-context conflict; persist and compare a hash of the full reviewed case/result, not only evidence.sha256.
Useful? React with 👍 / 👎.
Purpose
Publish the recovered BackIntel demo proof of concept. Users can replay the recorded business demonstration and exercise the analysis workspace with clearly labeled fixtures.
Behavior and controls
tabellio.validation.json. Use the owner-approvedtabellio.demo.validation.jsonfor demo merge.Validation
Current candidate:
631f31b339071f2fff177bcdda0a579b83871fd6.Local affected analysis and campaign regressions passed, including consent revocation, late publication changes, active-job replay, source-spec changes, stale candidates, and correction lock reuse. Package integrity and existing playback tests passed. The cleaned ZIP passed a fresh stdlib-only installation, playback, review-save/reload, and immutable-file verification. The analysis UI build and dead-code check passed. The dependency graph was refreshed and relevant callers checked against current source.
The hosted browser simulation now checks that model selection shows calibration metrics and reveals held-out metrics after promotion. Hosted checks and native Codex code review must pass on this candidate before merge.
Local evidence:
~/Library/Application Support/BackIntel/Evidence/FinalNativeReview20261007/, particularlyActiveRunReplayLocks/,ConsentFreshnessFixture/, andPrivateGrantPackage/.The separate security re-review is unavailable because its paying account reached its Codex usage limit. The earlier PostgreSQL finding was repaired and verified by the required hosted authentication check.
Demo and migration boundary
No paid provider calls, deployment, or release publication are included. Fixtures do not establish live-model quality, autonomous data preparation, or measured business value. Actual Jev, CatBoost, and TabICLv2 acceptance remains governed by the separate full-product contract. OlistV2 remains removed.
Source specifications require fresh acknowledgment when first fingerprinted or changed. Previous answers become stale after implementation/configuration changes. Existing services, data volumes, and recovery work are preserved. Before recreating old containers, configure matching database passwords and the capability API token as documented; repository fixes do not secure already-running containers. Older predictor packages must be prepared again under the current implementation.
For the historical recording, extract
docs/demo/Packages/BackIntelDemoBusinessV4.zipinto a new folder. FromBackIntelDemo, runpython3 RunDemo.py --check, thenpython3 RunDemo.py. Python 3.10+ is sufficient; no model keys or database server are needed.