Skip to content

feat(pack): add optional shared browser runtime - #3396

Merged
fireairforce merged 2 commits into
nextfrom
zoomdong-pack-shared-runtime
Sep 29, 2026
Merged

fireairforce merged 2 commits into
nextfrom
zoomdong-pack-shared-runtime

Conversation

@fireairforce

Copy link
Copy Markdown
Member

Summary

Production browser entries currently embed a separate copy of the runtime. Add optimization.sharedRuntime: true to emit one reusable runtime asset while retaining each entry's bootstrap script. HTML generation and webpack-compatible stats include the runtime in the correct entry asset order. The option defaults to false and does not affect development, Node.js, or library builds; types, schema, documentation, and integration tests are included.

In a local 20-entry build of examples/multi-entries-heavy, the sum of gzip-compressed JavaScript assets decreased by 75.6 KiB (9.9%). A single page loaded one additional script and 110 more gzip bytes, so this is opt-in for applications that benefit from sharing across entries. These measurements do not establish a build-time speedup.

Depends on utooland/next.js#197. This PR advances the submodule to its c037804473bb3ed17b47174258549f12a80854d8 commit; merge the fork prerequisite first. The short-session change in #3395 is independent.

Test Plan

  • Native addon and schema generator built with cargo build --locked --profile release-local -p pack-napi -p pack-schema --features plugin; generated schema and compiled both Pack and Pack Shared TypeScript.
  • Seven tests passed across sharedRuntime.test.ts, serveStats.test.ts, and serveClientPaths.test.ts, including actual builds, generated HTML ordering, runtime stats, unchanged development/Node output, and Node dynamic import execution.
  • Default runtime snapshot passed: cargo test --locked --profile release-local -p pack-tests app_build_runtime -- --nocapture (one matching test; no snapshots updated).
  • Local Chrome validation passed all 26 harness commands for the formal configuration: dynamic imports, workers, root/non-root public paths, multiple entries, blocking/deferred HTML scripts, and independent bundles with distinct chunk-loading globals. The browser harness is local validation tooling and is not included in this PR.
  • cargo fmt, cargo clippy --all-targets -- -D warnings --no-deps, npx tombi format --check, npx biome ci, typos, and git diff --check passed.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T03:41:58.040886Z ccee183 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

Copy link
Copy Markdown

📊 Performance Benchmark Report (with-antd)

Utoopack Performance Report

Report ID: utoopack_performance_report_20260929_040050
Generated: 2026-09-29 04:00:50
Trace File: trace_antd.json (0.3GB, 0.82M spans)
Test Project: examples/with-antd


Executive Summary

Metric Value Assessment
Total Wall Time 6,600.7 ms Baseline
Total Thread Work (de-duped) 20,471.0 ms Non-overlapping busy time
Effective Parallelism 3.1x thread_work / wall_time
Working Threads 10 Threads with actual spans
Thread Utilization 31.0% ⚠️ Suboptimal
Total Spans 824,873 All B/E + X events
Meaningful Spans (>= 10us) 224,746 (27.2% of total)
Tracing Noise (< 10us) 600,127 (72.8% of total)

Build Phase Timeline

Shows when each build phase is active and how much CPU it consumes.
Self-Time is the time spent exclusively in that phase (excluding children).

Phase Spans Inclusive (ms) Self-Time (ms) Wall Range (ms)
Resolve 50,153 6,989.4 1,935.9 3,307.9
Parse 8,170 1,354.9 1,016.1 5,908.4
Analyze 145,083 44,392.2 9,502.1 5,825.3
Chunk 5,461 6,029.8 835.3 2,290.3
Codegen 12,525 2,558.1 1,541.3 1,978.3
Emit 32 37.6 18.8 8.0
Other 3,322 7,215.8 4,038.7 6,600.7

Workload Distribution by Diagnostic Tier

Category Spans Inclusive (ms) % Work Self-Time (ms) % Self
P0: Scheduling & Resolution 195,819 51,896.2 253.5% 11,623.3 56.8%
P1: I/O & Heavy Tasks 2,900 107.9 0.5% 89.1 0.4%
P2: Architecture (Locks/Memory) 0 0.0 0.0% 0.0 0.0%
P3: Asset Pipeline 24,808 9,987.4 48.8% 3,415.2 16.7%
P4: Bridge/Interop 0 0.0 0.0% 0.0 0.0%
Other 1,219 6,586.4 32.2% 3,760.6 18.4%

Top 20 Tasks by Self-Time

Self-time is the exclusive duration: time spent in the task itself, not in sub-tasks.
This is the most accurate indicator of where CPU cycles are actually spent.

Self (ms) Inclusive (ms) Count Avg Self (us) P95 Self (ms) Max Self (ms) % Work Task Name Top Caller
4,820.6 26,596.6 97,626 49.4 0.1 11.9 23.5% module module (62%)
2,439.4 2,761.0 19 128389.9 327.9 463.9 11.9% save snapshot persist (5%)
2,293.5 2,447.1 2,469 928.9 2.8 300.4 11.2% analyze ecmascript module module (71%)
1,528.4 14,229.5 34,950 43.7 0.1 6.1 7.5% process module process module (82%)
1,253.3 3,358.9 27,375 45.8 0.1 6.0 6.1% internal resolving internal resolving (77%)
955.8 1,294.6 6,003 159.2 0.6 40.0 4.7% parse ecmascript parse ecmascript (65%)
828.0 915.2 10,227 81.0 0.3 8.7 4.0% precompute code generation generate merged code (39%)
724.0 5,729.2 4,019 180.2 0.2 79.7 3.5% chunking chunking (47%)
717.9 847.1 7,324 98.0 0.4 102.2 3.5% compute async module info compute async module info (57%)
672.7 3,620.6 22,077 30.5 0.0 5.2 3.3% resolving module (55%)
625.1 2,057.6 944 662.2 1.6 178.0 3.1% generate merged code chunking (45%)
422.0 422.0 329 1282.8 1.1 237.8 2.1% generate source map code generation (83%)
345.9 750.9 131 2640.1 5.0 195.7 1.7% emit code generate merged code (41%)
322.8 322.8 9 35864.8 162.3 188.6 1.6% blocking save snapshot (67%)
291.3 1,220.9 1,969 148.0 0.2 66.8 1.4% code generation code generation (83%)
258.9 541.0 1,701 152.2 0.1 129.3 1.3% write all entrypoints to disk write all entrypoints to disk (17%)
107.4 296.3 1,371 78.3 0.1 9.4 0.5% compute async chunks compute async chunks (47%)
82.9 104.9 819 101.2 0.0 26.8 0.4% compute binding usage info compute binding usage info (48%)
60.3 60.3 2,165 27.9 0.0 3.1 0.3% read file parse ecmascript (91%)
45.5 80.6 1,870 24.3 0.0 16.1 0.2% collect mergeable modules collect mergeable modules (100%)

Critical Path Analysis

The longest sequential dependency chains that determine wall-clock time.
Focus on reducing the depth of these chains to improve parallelism.

Rank Self-Time (ms) Depth Path
1 652.5 3 persist → save snapshot → blocking
2 377.3 6 chunking → generate merged code → emit code → emit code → emit code → read file
3 321.8 2 save snapshot → blocking
4 300.8 10 module → module → module → ... → process module → analyze ecmascript module → analyze ecmascript module
5 285.0 3 process module → process module → analyze ecmascript module

Batching Candidates

High-volume tasks dominated by a single parent. If the parent can batch them,
it drastically reduces scheduler overhead.

Task Name Count Top Caller (Attribution) Avg Self P95 Self Total Self
process module 34,950 process module (82%) 43.7 us 0.07 ms 1,528.4 ms
internal resolving 27,375 internal resolving (77%) 45.8 us 0.09 ms 1,253.3 ms

Duration Distribution

Range Count Percentage
<10us 600,127 72.8%
10us-100us 146,665 17.8%
100us-1ms 66,718 8.1%
1ms-10ms 11,098 1.3%
10ms-100ms 225 0.0%
>100ms 40 0.0%

Action Items

  1. [P0] Focus on tasks with the highest Self-Time — these are where CPU cycles are actually spent.
  2. [P0] Use Batching Candidates to identify callers that should use try_join or reduce #[turbo_tasks::function] granularity.
  3. [P1] Check Build Phase Timeline for phases with disproportionate wall range vs. self-time (= serialization).
  4. [P1] Inspect P95 Self (ms) for heavy monolith tasks. Focus on long-tail outliers, not averages.
  5. [P1] Review Critical Paths — reducing the longest chain depth directly improves wall-clock time.
  6. [P2] If Thread Utilization < 60%, investigate scheduling gaps (lock contention or deep dependency chains).

Report generated by Utoopack Performance Analysis Agent

@fireairforce
fireairforce merged commit 434c0a5 into next Sep 29, 2026
45 checks passed
@fireairforce
fireairforce deleted the zoomdong-pack-shared-runtime branch September 29, 2026 09:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants