Skip to content

Tracking issue for chunks_exact/_mut; slice chunks with exact size #47115

Description

@sdroege

This is inspired by ndarray and generally seems to allow llvm to remove more bounds checks in the code using the iterator (because the slices will always be exactly the requested size), and doesn't require the caller to add additional checks.

A PR adding these for further discussion will come in a bit.

Open questions:

  • Should the new iterators panic if the slice is not divisible by the chunk_size, or omit any leftover elements.
    The latter is implemented right now and very similar to how zip works and @shepmaster even argues that without this, this iterator is kind of useless and the optimization should be implemented as part of the normal chunks iterator (which seems non-trivial, see ).
    Omission of leftover elements is also how this iterator is implemented in ndarray (but far more general).
A function for getting access to the remainder exists on the iterator, similar to how `slice::Iter` and `slice::IterMut` give access to the tail (note: the remainder are the odd elements that don't completely fill a chunk, it's not the tail!)
**The majority of people (who spoke up here) seem to prefer the non-panicking behaviour**

- [x] Should it be called `exact_chunks` or `chunks_exact`? The former is how it's called in `ndarray`, the latter is [potentially more discoverable](https://github.com/rust-lang/rust/issues/47115#issuecomment-403090815) in e.g. IDEs.
**It was renamed to `chunks_exact`**

Activity

  1. sdroege commented on Jan 2, 2018

    @sdroege
    ContributorAuthor

    Example can be found here

    The relevant part with differences in the assembly is

    before:

    .LBB4_24:
      cmp r11, 4
      mov eax, 4
      cmovb rax, r11
      test rbx, rbx
      je .LBB4_18
      cmp rbx, 4
      mov edx, 4
      cmovb rdx, rbx
      test r13, r13
      je .LBB4_18
      mov qword ptr [rbp - 96], rax
      mov qword ptr [rbp - 48], rsi
      mov qword ptr [rbp - 56], r9
      cmp r11, 3
      jbe .LBB4_27
      mov qword ptr [rbp - 96], rdx
      mov qword ptr [rbp - 48], rsi
      mov qword ptr [rbp - 56], r9
      cmp rbx, 3
      jbe .LBB4_29
      cmp rax, 1
      je .LBB4_39
      lea r10, [r13 + rax]
      sub r11, rax
      lea r12, [r15 + rdx]
      sub rbx, rdx
      cmp rax, 3
      jb .LBB4_41
      je .LBB4_42
      movzx r14d, byte ptr [r13]
      movzx r8d, byte ptr [r13 + 1]
      movzx eax, byte ptr [r13 + 2]
      imul r13d, eax, 19595
      imul edi, r8d, 38470
      imul eax, r14d, 7471
      add eax, edi
      add eax, r13d
      shr eax, 16
      mov byte ptr [r15], al
      cmp rdx, 1
      je .LBB4_44
      mov byte ptr [r15 + 1], al
      cmp rdx, 3
      jb .LBB4_45
      mov byte ptr [r15 + 2], al
      je .LBB4_46
      mov byte ptr [r15 + 3], 0
      test r11, r11
      mov r15, r12
      mov r13, r10
      jne .LBB4_24

    after:

    .LBB5_18:
      test rsi, rsi
      je .LBB5_20
      add rdx, -4
      movzx r10d, byte ptr [rsi]
      movzx eax, byte ptr [rsi + 1]
      movzx ebx, byte ptr [rsi + 2]
      lea rsi, [rsi + 4]
      imul r13d, ebx, 19595
      imul eax, eax, 38470
      imul ebx, r10d, 7471
      add ebx, eax
      add ebx, r13d
      shr ebx, 16
      mov byte ptr [rcx], bl
      mov byte ptr [rcx + 1], bl
      mov byte ptr [rcx + 2], bl
      mov byte ptr [rcx + 3], 0
      lea rcx, [rcx + 4]
      cmp rdx, 4
      jae .LBB5_18
  2. bluss commented on Jan 2, 2018

    @bluss
    Contributor

    I don't want to derail your discussion too much. Const generics and value-level chunks both have their uses. I'm reminded of this existing implementation of the "const" kind of chunking, in this case in an iterator that actually allows access to the whole blocks and then the uneven tail at the end: BlockedIter. Note that a Block<Item=T> is an array of T.

  3. sdroege commented on Jan 2, 2018

    @sdroege
    ContributorAuthor

    Interesting, thanks for mentioning that. For my use case that would probably work more or less the same way, but it's slightly different indeed.

  4. bluss commented on Jan 2, 2018

    @bluss
    Contributor

    BlockedIter was developed while looking at exactly the hand off between the blocks and the elementwise tail; the idea was to avoid some of the loss that otherwise shows up in code that converts between slices and slice iterators. In this case it's the same pointer being bumped through the whole iteration.

  5. bluss commented on Jan 2, 2018

    @bluss
    Contributor

    The chunks iterators are a good candidate for zip specialization (TrustedRandomAccess trait)

  6. sdroege commented on Jan 2, 2018

    @sdroege
    ContributorAuthor

    True. I'll add that in a bit, as a separate PR for the existing chunked iterators and as a separate commit for the new ones.

  7. sdroege commented on Jan 2, 2018

    @sdroege
    ContributorAuthor

    I forgot to add some benchmark results earlier. This is with the code from #47115 (comment) and running on a 1920*1080*4 byte slice. Basically 2.46x as fast.

    running 2 tests
    test tests::bench_with_chunks       ... bench:   7,702,902 ns/iter (+/- 177,747)
    test tests::bench_with_exact_chunks ... bench:   3,132,468 ns/iter (+/- 202,032)
    
  8. added a commit that references this issue on Jan 4, 2018
  9. added a commit that references this issue on Jan 5, 2018
  10. 71 remaining items

  11. removed
    proposed-final-comment-periodProposed to merge/close by relevant subteam, see T-<team> label. Will enter FCP once signed off.
    on Oct 17, 2018
  12. rfcbot commented on Oct 17, 2018

    @rfcbot

    🔔 This is now entering its final comment period, as per the review above. 🔔

  13. alexcrichton commented on Oct 18, 2018

    @alexcrichton
    Member

    Ok! It's been quite awhile here so I think it's ok to shirt circuit the FCP slightly, @sdroege want to send the stabilization PR?

  14. SimonSapin commented on Oct 18, 2018

    @SimonSapin
    Contributor

    Should we consider #54580 to be "trivially enough" similar to this to stabilize at the same time?

  15. sdroege commented on Oct 18, 2018

    @sdroege
    ContributorAuthor

    @sdroege want to send the stabilization PR?

    Yeah, preparing a PR now. I'll add rchunks in a separate commit into the same PR if it's agreed that it should be part of that. But I guess first of all a review of #54580 should be done.

  16. sdroege commented on Oct 18, 2018

    @sdroege
    ContributorAuthor

    There's now #55178 for the stabilization of this here (but not rchunks).

  17. added a commit that references this issue on Oct 18, 2018
  18. added and removed
    final-comment-periodIn the final comment period and will be merged soon unless new substantive objections are raised.
    on Oct 27, 2018
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    B-unstableBlocker: Implemented in the nightly compiler and unstable.C-tracking-issueCategory: An issue tracking the progress of sth. like the implementation of an RFCT-libs-api[DEPRECATED; DO NOT USE]disposition-mergeThis issue / PR is in PFCP or FCP with a disposition to merge it.finished-final-comment-periodThe final comment period is finished for this PR / Issue.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions