Skip to content

Add merge yields pass - #162429

Open
diondokter wants to merge 12 commits into
rust-lang:mainfrom
diondokter:collapse_yields
Open

diondokter wants to merge 12 commits into
rust-lang:mainfrom
diondokter:collapse_yields

Conversation

@diondokter

@diondokter diondokter commented Sep 7, 2026 •

Copy link
Copy Markdown
Member

Part of Async statemachine optimisation project goal

r? dingxiangfei2009

Added a pass that merges functionally identical yields right before the async StateTransform pass. The StateTransform pass creates a unique state for every yield, so by merging yields we reduce the amount of states in the state machines.

See the module documentation for more information.

As for how much binary size is saved, that's hard to say in general. There's plenty of code that doesn't have identical yields. But when there is, the savings can be huge. In this project it saves 616 bytes. For a customer, when they applied this pass manually, it saved ~2kb in one instance.

It'll also play nice with the next optimization I'm going to work on, where we optimize async functions with a single await.

I think the biggest risk in the PR are the compare functions for the basic blocks and their parts. That's most of the code in the pass and can break if the basic block structures are updated but this code is forgotten. I don't know how to solve/improve that, so I'm curious if anyone can think of something. The threat is that the compare functions say the blocks are identical, but they're not in reality.

Also, the pass is currently always enabled. I don't know what people want there. Does it need an unstable flag? Or depend on the mir opt level?

@rustbot

rustbot commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator

Some changes occurred to MIR optimizations

cc @rust-lang/wg-mir-opt

@rustbot rustbot added S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. labels Sep 7, 2026
@rustbot

This comment has been minimized.

@rustbot rustbot added has-merge-commits PR has merge commits, merge with caution. S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. labels Sep 7, 2026
@rustbot

This comment has been minimized.

@rustbot rustbot removed has-merge-commits PR has merge commits, merge with caution. S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. labels Sep 7, 2026
@diondokter diondokter changed the title Add merge yields Add merge yields pass Sep 7, 2026
@rust-log-analyzer

This comment has been minimized.

@saethlin

Copy link
Copy Markdown
Member

I think there are soundness flaws in this MIR transform. This program when run with Miri reports UB with the new pass:

use std::future::Future;
use std::pin::{Pin, pin};
use std::task::{Context, Poll, Waker};

struct YieldOnce(bool);
impl Future for YieldOnce {
    type Output = (); 
    fn poll(mut self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<()> {
        if self.0 {
            Poll::Ready(())
        } else {
            self.0 = true;
            Poll::Pending
        }   
    }   
}

fn yield_once() -> YieldOnce {
    YieldOnce(false)
}

fn block_on<F: Future>(f: F) -> F::Output {
    let mut f = pin!(f);
    let mut cx = Context::from_waker(Waker::noop());
    loop {
        if let Poll::Ready(v) = f.as_mut().poll(&mut cx) {
            return v;
        }   
    }   
}

async fn a(arg: i32) -> i32 {
    yield_once().await;
    arg 
}

async fn b(val: bool) -> i32 {
    if val { a(1).await } else { a(-1).await }
}

fn main() {
    block_on(b(true));
}

@rustbot

rustbot commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

This PR was rebased onto a different main commit. Here's a range-diff highlighting what actually changed.

Rebasing is a normal part of keeping PRs up to date, so no action is needed—this note is just to help reviewers.

@rust-log-analyzer

This comment has been minimized.

@diondokter

Copy link
Copy Markdown
Member Author

Alright, the miri issue should be resolved now @saethlin
Also good to know that running the test suite doesn't mean these types of issue are checked.

The problem was that we were moving locals before they were initialized. Now that's being checked using MaybeInitializedPlaces and if the local isn't initialized then the move is omitted

@Amanieu

Amanieu commented Oct 1, 2026

Copy link
Copy Markdown
Member

I have serious concerns about this pass as it is currently written:

  • It checks every possible pair of Yield terminators, which is already quadratic. Then, for each pair it performs an unbounded scan of successor blocks, which can result in cubic total complexity (technically O(yields^2 * blocks)).
  • The matching logic is not handling a lot of cases:
    • It compares the base local of a place and ignores projections.
    • It ignores the values tested by SwitchInt.
    • It compares a traversal sequence without checking that corresponding terminator edges lead to corresponding blocks.
    • And others...
  • Inserting moves to translate between locals is not a valid implementation of local renaming because of borrowing. What you instead need is local unification, where all references to one local in a function are replaced with another local. To do this correctly you need a lifetime analysis that proves the two merged locals don't have overlapping allocation lifetimes.

I would recommend re-framing this as a basic block deduplication pass: if we can show that block A is identical to block B then we can redirect all control flow targeting block B to point to block A instead and eliminate block B. This achieves the desired effect of multiple Yield terminators sharing the same continuation blocks.

The implementation should be done in several stages.

Stage 1: Merge identical blocks only

At this stage, only blocks that have exactly the same statements and terminator can be merged. All fields and enum variants of the statements and terminators must be checked, as well as block-specific metadata like is_cleanup. There is some allowance for ignoring differences in debuginfo and spans, but this must be done carefully and explicitly.

Stage 2: Merge identical blocks with different successors

This extends stage 1 by merging blocks with different successors. This requires recursively proving that those successors are also identical. The current implementation using all_successors is wrong. Instead, you should use a worklist of BasicBlock pairs. For each pair, ensure that all statements are identical. Then ensure that the terminators are identical except for the successor blocks. Then check the successor blocks in each position: if they are the same or that pair has already been processed in the worklist, then it's fine. Otherwise push that pair to the worklist to continue proving that the pair of successor blocks are identical.

Once the worklist is empty, you have created a mapping of pairs of blocks that are known to be identical. You can now merge the pairs by redirecting all incoming control flow to a single block. Make sure to handle multiple groups of overlapping pairs correctly, I recommend using UnionFind for this.

Stage 3: Merge identical blocks with different locals

This gets more complicated, but is still doable. When 2 statements differ only in the locals they use, you need to prove that you can merge those 2 locals into a single local. This is only possible if the 2 locals have the same type and their allocation lifetimes do not overlap.

The PreciseLiveness analysis from #163337 gives you this information, although it is tied to the new MIR semantics from rust-lang/rfcs#3943. Have a look at #163338 to see how it is used, though in your case it will be simpler: you only need local-to-local mappings since projections are required to precisely match between the two blocks.

The basic idea is to keep track of pairs of locals that need to be unified as you go through blocks. Before adding a pair of locals you must check their type and that neither overlaps with the other or any other local that has previously been merged with them.

Once you have proved that you can unify the pairs of locals, then you've also proven that the blocks are identical, which makes it legal to merge them.

You will likely need to perform the same alias fixup and storage statement reconstruction as MoveElimination. The latter is particularly important since storage statements directly affect the set of locals that are saved across yield points.

Stage 4: Avoiding quadratic scanning

This is the part that I am less sure about, but at least it isn't necessary for correctness and can be deferred. The current algorithm scans the successors of every pair of yields, which ends up doing a lot of redundant work. It may instead be possible to iteratively merge common tail blocks by scanning backwards: unifying successors before predecessors allows block equivalences to expose further merge opportunities without repeatedly scanning sub-graphs of the function. At that point this becomes a general block deduplication optimization that is no longer specific to Yield. I'm not sure how this will handle loops, and will most likely require more thought.

@dingxiangfei2009

dingxiangfei2009 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

I would say that stage 1 and 2 are rarely going to be applicable due to how we lower .await expressions: if we see a <future expr>.await, the future will get pinned into a fresh local, always. These two options are untenable.

@dingxiangfei2009

dingxiangfei2009 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

But I have been thinking, whether we actually could do something during the MIR building phase.

For .await expressions, the future local is always moved, pinned onto a fresh local as of today, and from there the code starts a poll loop. This is a self-contained gadget in the AST2HIR lowering defined by fn lower_expr_await. I think we can do something about it, exposing the templating structure of such a gadget so that they get deduplicated again at THIR lowering.

@diondokter

Copy link
Copy Markdown
Member Author

Hey, thanks for the review!

This is the first time touching Rust's MIR, so it would've been amazing if I did it correct the first time already.

As for the complexity, I don't think that should be a huge problem since most functions don't have that many awaits. But I'll have to defer to the project on what's acceptable there.

Too bad inserting moves is not correct. I did consider doing local unification, but that seemed more difficult and so I chose the way I did.

I would say that stage 1 and 2 are rarely going to be applicable

I think these stages aren't meant to be that useful on their own, but to split the work in separately reviewable chunks.

So I think I will look into the proposed stages and eventually close this PR and open new ones for the stages.

@Amanieu

Amanieu commented Oct 2, 2026

Copy link
Copy Markdown
Member

For .await expressions, the future local is always moved, pinned onto a fresh local as of today, and from there the code starts a poll loop. This is a self-contained gadget in the AST2HIR lowering defined by fn lower_expr_await. I think we can do something about it, exposing the templating structure of such a gadget so that they get deduplicated again at THIR lowering.

Would it make sense to have a dedicated Await terminator in MIR that is later expanded into a loop?

@Amanieu

Amanieu commented Oct 2, 2026

Copy link
Copy Markdown
Member

I think these stages aren't meant to be that useful on their own, but to split the work in separately reviewable chunks.

So I think I will look into the proposed stages and eventually close this PR and open new ones for the stages.

It was more of a way to explain the optimization in a way that you can easily demonstrate its correctness. In terms of review, it's probably fine to do stage 1+2 together.

@dingxiangfei2009

dingxiangfei2009 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Await

Yes, @diondokter and I agreed and have believed that this would be more promising, but we took a more conservative approach which is to wrap the expression in a pseudo HIR node. That way the would be still freedom for us to maintain the expression lowering contained rather than dealing it at MIR runtime lowering: that would be much harder, I am afraid.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants