Repository navigation
Iterator Default Implementation for position() is slow #119551
Description
Activity
- addedneeds-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jan 3, 2024 The
try_fold-based implementation is likely intended to speed up position-finding inChainorFlatteniterators. So measuring range or slice iters doesn't exercise it well.That said, they should be equally fast for slices. Going through
try_foldshould end up with a very similar white-let loop.The call-tree is deeper, so this might affect optimizations. Have you tried compiling with 1CGU or LTO to see if that makes a difference? Benchmarks are fickle.
- addedI-slowIssue: Problems and improvements with respect to performance of generated code.Issue: Problems and improvements with respect to performance of generated code.A-iteratorsArea: IteratorsArea: IteratorsT-libsRelevant to the library team, which will review and decide on the PR/issue.Relevant to the library team, which will review and decide on the PR/issue.and removedneeds-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jan 4, 2024 The godbolt example does show that try_fold version results in a few more instructions.
The LLVM IR looks like it's trying to do some branchless version of the accumulator increments here:
%_30.i.i.i = getelementptr inbounds i32, ptr %_30.i12.i.i, i64 1, !dbg !71 %.val.i.i = load i32, ptr %_30.i12.i.i, align 4, !dbg !84, !noalias !85, !noundef !11 %_0.i.i.i.i.i = icmp ne i32 %.val.i.i, %0, !dbg !91 %_8.0.i.i.i.i = zext i1 %_0.i.i.i.i.i to i64, !dbg !103 %_0.sroa.3.0.i.i.i.i = add i64 %accum.0.i.i, %_8.0.i.i.i.i, !dbg !103while
%_30.i.i.i = getelementptr inbounds i32, ptr %_30.i.i56.i, i64 1, !dbg !81 %2 = add nuw nsw i64 %accum.07.i, 1, !dbg !94 %3 = icmp eq ptr %_30.i.i.i, %end_or_len, !dbg !36is an unconditional increment sitting in its own block.
Maybe ControlFlow is not well-suited for carrying the value and the loop state at the same time? CC @scottmcm
The
try_fold-based implementation is likely intended to speed up position-finding inChainorFlatteniterators. So measuring range or slice iters doesn't exercise it well.I did some quick measurements for a chain of vecs:
position fastest │ slowest │ median │ mean │ samples │ iters ├─ chain_new 477 µs │ 651.6 µs │ 487.3 µs │ 491.9 µs │ 204 │ 204 ╰─ chain_old 361.1 µs │ 651.4 µs │ 377 µs │ 380.2 µs │ 263 │ 263So the current impl does indeed work better for chains
That said, they should be equally fast for slices. Going through
try_foldshould end up with a very similar white-let loop.Yes, the difference is the ControlFlow abstraction
The call-tree is deeper, so this might affect optimizations. Have you tried compiling with 1CGU or LTO to see if that makes a difference? Benchmarks are fickle.
Yes:
[profile.bench] incremental = false codegen-units = 1 lto ="fat"No real differences though, results were with these settings.
Maybe ControlFlow is not well-suited for carrying the value and the loop state at the same time? CC @scottmcm
This made me retest an alternative I used previously (reducing ControlFlow size from <usize, usize> to <usize, ()>):
let mut accum = 0; self.try_fold((), |_, x| if predicate(x) { ControlFlow::Break(accum) } else { accum += 1; ControlFlow::Continue(()) }).break_value()And this one actually beats the current one for the chain test:
position fastest │ slowest │ median │ mean │ samples │ iters ├─ chain_alt 264.9 µs │ 450.4 µs │ 271.3 µs │ 276.9 µs │ 361 │ 361 ├─ chain_new 314.5 µs │ 413.3 µs │ 322.3 µs │ 326.4 µs │ 307 │ 307 ╰─ chain_old 361.6 µs │ 526.2 µs │ 371 µs │ 373.4 µs │ 268 │ 268The chain_alt is even faster, but it's an attempt to keep the current style, updated for keeping the counter outside the fold:
#[inline] fn check<'a, T>( mut predicate: impl FnMut(T) -> bool + 'a, acc: &'a mut usize, ) -> impl FnMut((), T) -> ControlFlow<usize, ()> + 'a { move |_, x| { if predicate(x) { ControlFlow::Break(*acc) } else { *acc += 1; ControlFlow::Continue(()) } } } let mut acc = 0; self.try_fold((), check(predicate, &mut acc)).break_value()It isn't pretty, but it works 😅
Godbolt for the alt version with chains
The Assembly and IR are too long for me too really identify where the benefit comes from, curiously the faster version has more instructions, but I guess with all the jumps everywhere that doesn't mean a lot.I am interested in testing these more thoroughly, is there a nice resource with some real world data to test them on?
Edit: Went digging in recent projects and I had a solution for Advent of Code Day 6 using .rposition() and .position() with the relevant part being:
let mut hold_durations = (0..time).map(|hold| hold * (time - hold)); hold_durations.rposition(|d| d > distance).unwrap() - hold_durations.position(|d| d > distance).unwrap() + 1And changing to the alt version (with check(predicate, acc)) improved runtime from 8.8 ms to 6.2 ms, a 40% speedup! (I know there is an algebraic solution to the problem, but still)
I am interested in testing these more thoroughly, is there a nice resource with some real world data to test them on?
You could make a PR and we can do a perf run. It doesn't have dedicated benchmarks that stress
positionand I haven't checked if any of the runtime benchmarks make use ofpositionat all, but the compiler also makes use of it internally so it might show up incheckbuilds too.- added a commit that references this issue
on Jan 5, 2024 - added a commit that references this issue
on Jan 5, 2024 - added a commit that references this issue
on Jan 5, 2024 - linked a pull request that will close this issueRewrite Iterator::position default impl #119599
on Jan 5, 2024 @the8472 Yes, we learned in #76746 (comment) that if you don't need move access to the accumulator -- either because
&mutis fine, as in that thread, or seemingly because you canCopy, as seen here -- then it's better to usetry_for_eachwith aFnMutclosure instead of needing LLVM to understand the threading through thetry_fold+Fn.(Maybe this'll be better once LLVM does better with 2-variant enums and stops losing information about them, but for now we're stuck with this.)
Reacted by the8472 and marthadev
The compiler struggles to optimize the current default implementation for position on Iterators.
Simplifying the implementation has increased the efficiency in various scenarios OMM:
Old position():
Simplified version (There are several alternative implementations producing similar results as well):
Bench results:
bench.rs
Benchmarks are named after the iterator and were run using Divan.
Ranges benefit massively (big ones by multiple orders of magnitude), while for other types, speedups are usually between 0% and 100%.
A specific implementation for range calculating the position would of course be even faster, but I don't think position() is called that often on ranges.
The implementation is similar to the slice::iter::Iter one, which was introduced to reduce compile times #72166, so perhaps this would help in that regard as well.
However, the compiler appears to occasionally forget how to optimize the function after editing unrelated code. This also happens with the current implementation, so I don't know how reproducible the speedups are (FWIW Ranges have never failed to improve, though).
I did not attempt to benchmark compilation speed, but according to godbolt the memory usage is down from 11.8 MB to 8.3 MB which looks promising.
I was motivated by a thread on reddit, where some context and speculation can be found.