Skip to content

arc: improve codegen of drop - #162178

Open
ruriww wants to merge 4 commits into
rust-lang:mainfrom
ruriww:main
Open

ruriww wants to merge 4 commits into
rust-lang:mainfrom
ruriww:main

Conversation

@ruriww

@ruriww ruriww commented Sep 2, 2026 •

Copy link
Copy Markdown

View all comments

Move the fence into drop_slow, which preserves the correct behavior and moving an actual instruction on arches like riscv to the outlined path.

@rustbot rustbot added S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. T-libs Relevant to the library team, which will review and decide on the PR/issue. labels Sep 2, 2026
@ruriww

ruriww commented Sep 2, 2026

Copy link
Copy Markdown
Author
pub fn drop_it_like_its_hot(_: Arc<()>) {}
 	sd	a0, 0(sp)
 	amoadd.d.rl	a0, a1, (a0)
 	li	a1, 1
-	bne	a0, a1, .LBB6_2
-	fence	r, rw
+	beq	a0, a1, .LBB6_2
+	ld	ra, 8(sp)
+	addi	sp, sp, 16
+	ret
+.LBB6_2:
 	mv	a0, sp
 	call	Arc::drop_slow
-.LBB6_2:
 	ld	ra, 8(sp)
 	addi	sp, sp, 16
 	ret

@ruriww
ruriww marked this pull request as ready for review September 2, 2026 11:03
@rustbot rustbot added the S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. label Sep 2, 2026
@rustbot rustbot removed the S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. label Sep 2, 2026
@rustbot

rustbot commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

Thanks for the pull request, and welcome! The Rust Project has assigned @Mark-Simulacrum (or someone else) to review your changes, you should hear from them (or someone else) within the next two weeks.

Please see the contribution instructions and our LLM policy for more information.

Why was this reviewer chosen?

The reviewer was selected based on:

  • Owners of files modified in this PR: libs
  • libs expanded to 12 candidates
  • Random selection from JohnTitor, Mark-Simulacrum, clarfonthey

@Kobzol

Kobzol commented Sep 2, 2026

Copy link
Copy Markdown
Member

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 2, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Sep 2, 2026
arc: improve codegen of drop
@rust-bors

rust-bors Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

☀️ Try build successful (CI)
Build commit: 7aa9fe8 (7aa9fe8c23689bedc5cb35a311d46262349ee5f5)
Base parent: 59dabe5 (59dabe56f7b78d9dd427645cd7f8e7fd426724cc)

@rust-timer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (7aa9fe8): comparison URL.

Overall result: ❌✅ regressions and improvements - BENCHMARK(S) FAILED

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

❗ ❗ ❗ ❗ ❗
Warning ⚠️: The following benchmark(s) failed to build:

  • Job failure

❗ ❗ ❗ ❗ ❗

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
0.4% [0.2%, 0.5%] 11
Improvements ✅
(primary)
-0.5% [-0.5%, -0.5%] 2
Improvements ✅
(secondary)
-0.5% [-0.5%, -0.5%] 1
All ❌✅ (primary) -0.5% [-0.5%, -0.5%] 2

Max RSS (memory usage)

Results (primary -1.3%, secondary -3.4%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
2.8% [1.7%, 3.9%] 2
Regressions ❌
(secondary)
- - 0
Improvements ✅
(primary)
-4.0% [-6.2%, -2.4%] 3
Improvements ✅
(secondary)
-3.4% [-3.7%, -3.0%] 2
All ❌✅ (primary) -1.3% [-6.2%, 3.9%] 5

Cycles

Results (primary -2.5%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
- - 0
Improvements ✅
(primary)
-2.5% [-2.8%, -2.3%] 2
Improvements ✅
(secondary)
- - 0
All ❌✅ (primary) -2.5% [-2.8%, -2.3%] 2

Binary size

Results (primary 0.2%, secondary -0.6%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
0.3% [0.0%, 0.8%] 16
Regressions ❌
(secondary)
0.1% [0.1%, 0.1%] 1
Improvements ✅
(primary)
-0.0% [-0.1%, -0.0%] 5
Improvements ✅
(secondary)
-0.9% [-1.7%, -0.0%] 2
All ❌✅ (primary) 0.2% [-0.1%, 0.8%] 21

Bootstrap: missing data
Artifact size: 400.83 MiB -> 400.77 MiB (-0.01%)

@rustbot rustbot added perf-regression Performance regression. and removed S-waiting-on-perf Status: Waiting on a perf run to be completed. labels Sep 2, 2026
@ruriww

ruriww commented Sep 2, 2026

Copy link
Copy Markdown
Author

BENCHMARK(S) FAILED

Is this something I have to take care of? The instruction regression makes sense, it's demonstrated in the diff snippet I showed above.

@Kobzol

Kobzol commented Sep 2, 2026

Copy link
Copy Markdown
Member

No, you can ignore that, we have some intermittent trouble with git, sorry.

Comment thread library/alloc/src/sync.rs Outdated
Comment on lines +2970 to +2971
// This function was moved locally since there is only one caller,
// and makes it easier to reason about the the outlined fence.

@vilgotf vilgotf Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Drive-by review: this code comment looks like a commit message. I don't think it holds value by itself.

View changes since the review

@panstromek

Copy link
Copy Markdown
Contributor

tt-muncher regressions are probably real. It's a macro stress test and macro expansion uses Arc quite a bit.

@ruriww

ruriww commented Sep 3, 2026 •

Copy link
Copy Markdown
Author

tt-muncher regressions are probably real. It's a macro stress test and macro expansion uses Arc quite a bit.

I'm pretty certain this is caused by the cold_path addition. The move of the acquire line into drop_slow should be universally better though. Though I am curious to see how this holds up in user code. The branch hint could be worth keeping if it ends up being a net improvement. I guess the user benchmarks would be the runtime category, I wasn't quite familiar with the CI testing.

@rust-log-analyzer

This comment has been minimized.

@ruriww

ruriww commented Sep 4, 2026

Copy link
Copy Markdown
Author

Removed the cold hint. I had AI help me dig deeper into the performance numbers only, here are the results: PGO data shows that the drop_slow path is taken 68% of the time in the CI benchmarks, which makes sense, as many Arcs are only ever 1-counted. This would mean that user code that heavily refcounts things should use PGO to replicate what the cold_path addition would have done.

@Mark-Simulacrum Mark-Simulacrum left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Especially given the potential problems (see first comment), is there a concrete motivation for this change? The assembly diff in #162178 (comment) doesn't seem necessarily sufficient to motivate tweaking this code by itself.

View changes since this review

Comment thread library/alloc/src/sync.rs
Comment thread library/alloc/src/sync.rs
fn drop(&mut self) {
// Non-inlined part of `drop`.
#[inline(never)]
unsafe fn drop_slow<T: ?Sized, A: Allocator>(this: &mut Arc<T, A>) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why inline this into drop?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed a comment that was originally in the code stating that it was easier to reason about the fence that was removed from fn drop but added to fn drop_slow.

@Mark-Simulacrum

Copy link
Copy Markdown
Member

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rust-bors

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 13, 2026
rust-bors Bot pushed a commit that referenced this pull request Sep 13, 2026
arc: improve codegen of drop
@rust-bors

rust-bors Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

☀️ Try build successful (CI)
Build commit: 00d30ba (00d30ba0e863df00282674f152812cc5963072bf)
Base parent: 20d35a3 (20d35a3ae8f310f2a002e5f6e0bc583830010cd4)

@rust-timer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (00d30ba): comparison URL.

Overall result: ❌✅ regressions and improvements - please read:

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
0.2% [0.2%, 0.2%] 1
Regressions ❌
(secondary)
0.4% [0.4%, 0.5%] 10
Improvements ✅
(primary)
-0.4% [-0.6%, -0.3%] 2
Improvements ✅
(secondary)
-0.6% [-0.7%, -0.6%] 4
All ❌✅ (primary) -0.2% [-0.6%, 0.2%] 3

Max RSS (memory usage)

Results (primary 2.9%, secondary 0.8%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
2.9% [2.0%, 3.7%] 2
Regressions ❌
(secondary)
4.7% [3.5%, 5.3%] 4
Improvements ✅
(primary)
- - 0
Improvements ✅
(secondary)
-4.4% [-5.5%, -3.7%] 3
All ❌✅ (primary) 2.9% [2.0%, 3.7%] 2

Cycles

Results (secondary -1.1%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
4.1% [2.2%, 7.5%] 4
Improvements ✅
(primary)
- - 0
Improvements ✅
(secondary)
-5.3% [-10.1%, -2.8%] 5
All ❌✅ (primary) - - 0

Binary size

Results (primary 0.4%, secondary -0.8%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
0.4% [0.3%, 0.7%] 8
Regressions ❌
(secondary)
- - 0
Improvements ✅
(primary)
- - 0
Improvements ✅
(secondary)
-0.8% [-1.6%, -0.0%] 2
All ❌✅ (primary) 0.4% [0.3%, 0.7%] 8

Bootstrap: 497.229s -> 494.834s (-0.48%)
Artifact size: 406.94 MiB -> 406.90 MiB (-0.01%)

@rustbot rustbot removed the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 14, 2026
@ruriww

ruriww commented Sep 14, 2026

Copy link
Copy Markdown
Author

The same tt-muncher regression are odd, and I don't think that's to do with the changes here, x64 is implicitly AcqRel so the fence is a no-op.

@Mark-Simulacrum

Copy link
Copy Markdown
Member

@bors try @rust-timer queue

Let's get another perf run in to try to validate those numbers. It's possible that shuffling here has changed some inlining or other codegen decisions leading to slower compilation (even if the final assembly is the same).

Do you have any evidence that moving this instruction makes a practical difference, at least in a microbench of some kind?

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 27, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Sep 27, 2026
arc: improve codegen of drop
@rust-bors

rust-bors Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

☀️ Try build successful (CI)
Build commit: 12ea40b (12ea40b08120d0afc0a6442aaec4d559944d552f)
Base parent: 8b5b91f (8b5b91f7bba722673c989e2964c8a094ce3c7939)

@rust-timer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (12ea40b): comparison URL.

Overall result: ❌✅ regressions and improvements - please read:

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
0.4% [0.2%, 0.5%] 11
Improvements ✅
(primary)
-0.4% [-0.6%, -0.2%] 3
Improvements ✅
(secondary)
-0.4% [-0.6%, -0.3%] 3
All ❌✅ (primary) -0.4% [-0.6%, -0.2%] 3

Max RSS (memory usage)

Results (primary 0.7%, secondary 1.0%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
2.0% [1.2%, 2.4%] 3
Regressions ❌
(secondary)
1.8% [1.2%, 3.1%] 4
Improvements ✅
(primary)
-3.1% [-3.1%, -3.1%] 1
Improvements ✅
(secondary)
-2.2% [-2.2%, -2.2%] 1
All ❌✅ (primary) 0.7% [-3.1%, 2.4%] 4

Cycles

Results (primary -2.4%, secondary 0.1%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
3.2% [3.2%, 3.2%] 1
Improvements ✅
(primary)
-2.4% [-2.4%, -2.4%] 1
Improvements ✅
(secondary)
-1.4% [-2.1%, -0.7%] 2
All ❌✅ (primary) -2.4% [-2.4%, -2.4%] 1

Binary size

Results (primary 0.4%, secondary -0.8%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
0.4% [0.3%, 0.8%] 5
Regressions ❌
(secondary)
- - 0
Improvements ✅
(primary)
- - 0
Improvements ✅
(secondary)
-0.8% [-1.7%, -0.0%] 2
All ❌✅ (primary) 0.4% [0.3%, 0.8%] 5

Bootstrap: 488.457s -> 488.75s (0.06%)
Artifact size: 406.36 MiB -> 406.38 MiB (0.01%)

@rustbot rustbot removed the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 27, 2026
@Mark-Simulacrum

Mark-Simulacrum commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

It looks like the tt-muncher regression is real and is in Arc-related code... I think before we merge this investigating the details here to figure out where the extra instructions are coming from is warranted. Can you try to dig into that in more detail, e.g., compiling a few programs that drop Arcs in x86_64 code and see if you can find a case with extra instructions?

This is the top of the diff from cachegrind:

$ cargo build --release -p collector && ./target/release/collector profile_local cachegrind +8b5b91f7bba722673c989e2964c8a094ce3c7939 --rustc2 +12ea40b08120d0afc0a6442aaec4d559944d552f --exact-match tt-muncher --profiles Check --scenarios Full
--------------------------------------------------------------------------------
-- File:function summary
--------------------------------------------------------------------------------
  Ir_________  file:function

<  11,828,832  ???:
   11,900,242    <rustc_parse::parser::Parser>::parse_token_tree
  -11,119,410    <alloc::sync::Arc<alloc::vec::Vec<rustc_ast::tokenstream::TokenTree>>>::drop_slow
   11,119,410    <alloc::sync::Arc<_, _> as core::ops::drop::Drop>::drop::drop_slow::<alloc::vec::Vec<rustc_ast::tokenstream::TokenTree>, alloc::alloc::Global>
     -567,477    <alloc::sync::Arc<rustc_middle::traits::ObligationCauseCode>>::drop_slow
      567,477    <alloc::sync::Arc<_, _> as core::ops::drop::Drop>::drop::drop_slow::<rustc_middle::traits::ObligationCauseCode, alloc::alloc::Global>

@rust-bors

rust-bors Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

☔ The latest upstream changes (presumably #163472) made this pull request unmergeable. Please resolve the merge conflicts by rebasing.

@Mark-Simulacrum Mark-Simulacrum added S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. and removed S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. labels Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

perf-regression Performance regression. S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. T-libs Relevant to the library team, which will review and decide on the PR/issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants