Repository navigation
Tracking issue for chunks_exact/_mut; slice chunks with exact size #47115
Description
Activity
Example can be found here
The relevant part with differences in the assembly is
before:
.LBB4_24: cmp r11, 4 mov eax, 4 cmovb rax, r11 test rbx, rbx je .LBB4_18 cmp rbx, 4 mov edx, 4 cmovb rdx, rbx test r13, r13 je .LBB4_18 mov qword ptr [rbp - 96], rax mov qword ptr [rbp - 48], rsi mov qword ptr [rbp - 56], r9 cmp r11, 3 jbe .LBB4_27 mov qword ptr [rbp - 96], rdx mov qword ptr [rbp - 48], rsi mov qword ptr [rbp - 56], r9 cmp rbx, 3 jbe .LBB4_29 cmp rax, 1 je .LBB4_39 lea r10, [r13 + rax] sub r11, rax lea r12, [r15 + rdx] sub rbx, rdx cmp rax, 3 jb .LBB4_41 je .LBB4_42 movzx r14d, byte ptr [r13] movzx r8d, byte ptr [r13 + 1] movzx eax, byte ptr [r13 + 2] imul r13d, eax, 19595 imul edi, r8d, 38470 imul eax, r14d, 7471 add eax, edi add eax, r13d shr eax, 16 mov byte ptr [r15], al cmp rdx, 1 je .LBB4_44 mov byte ptr [r15 + 1], al cmp rdx, 3 jb .LBB4_45 mov byte ptr [r15 + 2], al je .LBB4_46 mov byte ptr [r15 + 3], 0 test r11, r11 mov r15, r12 mov r13, r10 jne .LBB4_24
after:
.LBB5_18: test rsi, rsi je .LBB5_20 add rdx, -4 movzx r10d, byte ptr [rsi] movzx eax, byte ptr [rsi + 1] movzx ebx, byte ptr [rsi + 2] lea rsi, [rsi + 4] imul r13d, ebx, 19595 imul eax, eax, 38470 imul ebx, r10d, 7471 add ebx, eax add ebx, r13d shr ebx, 16 mov byte ptr [rcx], bl mov byte ptr [rcx + 1], bl mov byte ptr [rcx + 2], bl mov byte ptr [rcx + 3], 0 lea rcx, [rcx + 4] cmp rdx, 4 jae .LBB5_18
Reacted by Clément Renault, Henri Sivonen and Gurinder SinghI don't want to derail your discussion too much. Const generics and value-level chunks both have their uses. I'm reminded of this existing implementation of the "const" kind of chunking, in this case in an iterator that actually allows access to the whole blocks and then the uneven tail at the end: BlockedIter. Note that a
Block<Item=T>is an array of T.Interesting, thanks for mentioning that. For my use case that would probably work more or less the same way, but it's slightly different indeed.
BlockedIter was developed while looking at exactly the hand off between the blocks and the elementwise tail; the idea was to avoid some of the loss that otherwise shows up in code that converts between slices and slice iterators. In this case it's the same pointer being bumped through the whole iteration.
The chunks iterators are a good candidate for zip specialization (TrustedRandomAccess trait)
True. I'll add that in a bit, as a separate PR for the existing chunked iterators and as a separate commit for the new ones.
I forgot to add some benchmark results earlier. This is with the code from #47115 (comment) and running on a 1920*1080*4 byte slice. Basically 2.46x as fast.
running 2 tests test tests::bench_with_chunks ... bench: 7,702,902 ns/iter (+/- 177,747) test tests::bench_with_exact_chunks ... bench: 3,132,468 ns/iter (+/- 202,032)Reacted by Clément Renault and Kornel- added a commit that references this issue
on Jan 5, 2018 71 remaining items
- removedproposed-final-comment-periodProposed to merge/close by relevant subteam, see T-<team> label. Will enter FCP once signed off.Proposed to merge/close by relevant subteam, see T-<team> label. Will enter FCP once signed off.
on Oct 17, 2018 🔔 This is now entering its final comment period, as per the review above. 🔔
Ok! It's been quite awhile here so I think it's ok to shirt circuit the FCP slightly, @sdroege want to send the stabilization PR?
Should we consider #54580 to be "trivially enough" similar to this to stabilize at the same time?
There's now #55178 for the stabilization of this here (but not
rchunks).- added a commit that references this issue
on Oct 18, 2018 - addedfinished-final-comment-periodThe final comment period is finished for this PR / Issue.The final comment period is finished for this PR / Issue.and removedfinal-comment-periodIn the final comment period and will be merged soon unless new substantive objections are raised.In the final comment period and will be merged soon unless new substantive objections are raised.
on Oct 27, 2018
This is inspired by ndarray and generally seems to allow llvm to remove more bounds checks in the code using the iterator (because the slices will always be exactly the requested size), and doesn't require the caller to add additional checks.
A PR adding these for further discussion will come in a bit.
Open questions:
chunk_size, or omit any leftover elements.The latter is implemented right now and very similar to how
zipworks and @shepmaster even argues that without this, this iterator is kind of useless and the optimization should be implemented as part of the normal chunks iterator (which seems non-trivial, see ).Omission of leftover elements is also how this iterator is implemented in ndarray (but far more general).