Skip to content

Vec<u8> clone in rustc 1.33.0 is 3 times slower than rustc 1.29.0 #57437

Description

@breezewish

Benchmark code:

#[bench]
fn bench(b: &mut test::Bencher) {
    let raw = vec![0u8; 1000];
    b.iter(|| {
        test::black_box(test::black_box(&raw).clone());
    });
}

In rustc 1.29.0-nightly (4f3c7a4 2018-07-17): 32 ns/iter (+/- 34)
In rustc 1.33.0-nightly (9eac386 2018-12-31): 127 ns/iter (+/- 45)

Activity

  1. killercup commented on Jan 8, 2019

    @killercup
    Contributor

    Can you try this? #47745 (comment)

  2. added
    I-slowIssue: Problems and improvements with respect to performance of generated code.
    on Jan 8, 2019
  3. changed the title [-]&[u8] clone in rustc 1.33.0 is 50% slower than rustc 1.29.0[/-] [+]Vec<u8> clone in rustc 1.33.0 is 3 times slower than rustc 1.29.0[/+] on Jan 8, 2019
  4. breezewish commented on Jan 8, 2019

    @breezewish
    Author

    @killercup

    Hi, I tried with the following profile:

    [profile.bench]
    lto = false
    opt-level = 3
    debug = true
    codegen-units = 1
    
    [profile.release]
    lto = false
    opt-level = 3
    debug = true
    codegen-units = 1

    and

    [profile.bench]
    lto = true
    opt-level = 3
    debug = true
    codegen-units = 1
    
    [profile.release]
    lto = true
    opt-level = 3
    debug = true
    codegen-units = 1

    the outcome is similar.

  5. ollie27 commented on Jan 8, 2019

    @ollie27
    Contributor

    I'd guess this is due to #55238. You could try using jemallocator to confirm.

  6. brson commented on Jan 8, 2019

    @brson
    Contributor

    @ollie27 Oh very good guess. That would explain a lot. I do think that @breeswish's benchmarks are running against system malloc in the 'after' run. We are running at least one set of benchmarks with jemalloc both before/after: https://gist.github.com/brson/13586d9f12f3af5c8377628c3d0f12d0#file-benchcmp-tikv and have seen regressions there too, but not investigated.

    We'll fix our side to make sure we are comparing jemalloc to jemalloc then see how our benchmarks look.

  7. brson commented on Jan 9, 2019

    @brson
    Contributor

    What I reported yesterday about not comparing allocator to allocator looks to be incorrect. @breeswish's benchmarks may have been using the same jemalloc. Still investigating.

  8. mati865 commented on Jan 10, 2019

    @mati865
    Member

    What is your system?

    Since switch system allocator I'm seeing small performance increase on 3 systems with glibc 2.28 (Arch Linux, Fedora and Ubuntu).

    With your benchmark I was getting results so close they weren't reliable.
    These are results with let raw = vec![0u8; 1000000];:

    $ cargo +nightly-2018-07-17 bench
    [...]
    test bench ... bench:      19,454 ns/iter (+/- 241)
    
    $ cargo +nightly-2018-07-17 bench
    [...]
    test bench ... bench:      19,422 ns/iter (+/- 207)
    
    $ cargo +nightly-2018-12-31 bench
    [...]
    test bench ... bench:      19,378 ns/iter (+/- 2,560)
    
    $ cargo +nightly-2018-12-31 bench
    [...]
    test bench ... bench:      19,374 ns/iter (+/- 422)
    
    $ cargo +nightly bench           
    [...]
    test bench ... bench:      19,352 ns/iter (+/- 7,552)
    
    $ cargo +nightly bench
    [...]
    test bench ... bench:      19,342 ns/iter (+/- 7,586)
    
  9. breezewish commented on Jan 10, 2019

    @breezewish
    Author

    Hi @mati865 My OS is MacOS 10.12.6. I will try again with jemalloc linked. During that, you may first view a result powered by Travis CI (although it may not be very stable, but still referable): https://travis-ci.com/breeswish/vec_clone_play

  10. mati865 commented on Jan 10, 2019

    @mati865
    Member

    @breeswish I don't use macOS so I cannot speak for it but for such old Linux distributions jemallocator should fix the performance.

  11. brson commented on Jan 23, 2019

    @brson
    Contributor

    After further investigation, there indeed wasn't a problem with Vec<u8>, so this can be closed.

  12. breezewish commented on Jan 23, 2019

    @breezewish
    Author

    I forced to use jemalloc and discovered that there is no notable difference in the case reported by this issue. So closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    I-slowIssue: Problems and improvements with respect to performance of generated code.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions