Skip to content

Optimize String.chunk with uniform fast-path and sub-binary slicing - #15959

Open
preciz wants to merge 1 commit into
elixir-lang:mainfrom
preciz:perf/string-chunk-reuse
Open

preciz wants to merge 1 commit into
elixir-lang:mainfrom
preciz:perf/string-chunk-reuse

Conversation

@preciz

@preciz preciz commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Assisted-by: Gemini:3.8-Flash
Assisted-by: Codex:GPT-6

Add fast-paths for fully valid and printable strings to return [string]
immediately, avoiding the chunk loop and allocations entirely for uniform
inputs.

In the fallback chunk loop:

  • Reuse the first decoded codepoint and classification, eliminating
    redundant decoding and the intermediate wrapper.
  • Replace repeated <> binary concatenation with zero-copy sub-binary
    slicing via binary_part/3, removing intermediate binary reallocations.

Bench:

# Run with: elixir bench.exs
Mix.install([:benchee])

path = "lib/elixir/lib/string.ex"

for {ref, module} <- [{"HEAD^1", Old}, {"HEAD", New}] do
  {source, 0} = System.cmd("git", ["show", "#{ref}:#{path}"], cd: __DIR__)

  source
  |> String.replace("defmodule String do", "defmodule #{inspect(module)} do")
  |> Code.compile_string(path)
end

Benchee.run(
  %{
    "HEAD^1" => fn {string, trait} -> Old.chunk(string, trait) end,
    "HEAD" => fn {string, trait} -> New.chunk(string, trait) end
  },
  inputs: %{
    "11B valid" => {"hello world", :valid},
    "11B printable" => {"hello world", :printable},
    "1200B valid" =>
      {String.duplicate("The quick brown fox jumps over the lazy dog! ✨ ", 25), :valid},
    "1200B printable" =>
      {String.duplicate("The quick brown fox jumps over the lazy dog! ✨ ", 25), :printable},
    "1200B mixed valid/invalid" =>
      {String.duplicate("hello" <> <<255, 254>> <> "world" <> <<128>>, 80), :valid},
    "1200B mixed printable" =>
      {String.duplicate("abc\x00def\tghi\x1b[0m", 80), :printable}
  },
  pre_check: :all_same,
  warmup: 1,
  time: 3,
  memory_time: 1
)

Results:

Operating System: Linux
CPU Information: AMD Ryzen 7 8845HS w
Number of Available Cores: 16
Available memory: 54.72 GB
Elixir 1.20.4
Erlang 29.0.5
JIT enabled: true

Benchmark suite executing with the following configuration:
warmup: 1 s
time: 3 s
memory time: 1 s
reduction time: 0 ns
parallel: 1
inputs: 11B printable, 11B valid, 1200B mixed printable, 1200B mixed valid/invalid, 1200B printable, 1200B valid
Estimated total run time: 1 min
Excluding outliers: false

##### With input 11B printable #####
Name             ips        average  deviation         median         99th %
HEAD         13.92 M       71.83 ns   ±544.10%          70 ns         170 ns
HEAD^1        1.22 M      819.24 ns   ±773.21%         712 ns        1724 ns

Comparison: 
HEAD         13.92 M
HEAD^1        1.22 M - 11.41x slower +747.41 ns

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1         2.30 KB - 42.14x memory usage +2.25 KB

**All measurements for memory usage were the same**

##### With input 11B valid #####
Name             ips        average  deviation         median         99th %
HEAD         17.70 M       56.50 ns  ±2338.01%          50 ns          80 ns
HEAD^1        1.31 M      764.56 ns  ±1064.34%         651 ns        1583 ns

Comparison: 
HEAD         17.70 M
HEAD^1        1.31 M - 13.53x slower +708.06 ns

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1         2.30 KB - 42.14x memory usage +2.25 KB

**All measurements for memory usage were the same**

##### With input 1200B mixed printable #####
Name             ips        average  deviation         median         99th %
HEAD         21.42 K       46.68 μs    ±32.67%       45.48 μs       56.20 μs
HEAD^1       13.02 K       76.82 μs    ±35.83%       72.26 μs      163.99 μs

Comparison: 
HEAD         21.42 K
HEAD^1       13.02 K - 1.65x slower +30.14 μs

Memory usage statistics:

Name      Memory usage
HEAD         202.66 KB
HEAD^1       245.34 KB - 1.21x memory usage +42.69 KB

**All measurements for memory usage were the same**

##### With input 1200B mixed valid/invalid #####
Name             ips        average  deviation         median         99th %
HEAD         24.21 K       41.30 μs    ±34.26%       40.17 μs       48.52 μs
HEAD^1       14.22 K       70.34 μs    ±42.25%       61.49 μs      168.72 μs

Comparison: 
HEAD         24.21 K
HEAD^1       14.22 K - 1.70x slower +29.04 μs

Memory usage statistics:

Name      Memory usage
HEAD         180.92 KB
HEAD^1       215.33 KB - 1.19x memory usage +34.41 KB

**All measurements for memory usage were the same**

##### With input 1200B printable #####
Name             ips        average  deviation         median         99th %
HEAD        215.93 K        4.63 μs    ±39.48%        4.48 μs        6.42 μs
HEAD^1       16.09 K       62.16 μs    ±34.26%       60.44 μs       73.20 μs

Comparison: 
HEAD        215.93 K
HEAD^1       16.09 K - 13.42x slower +57.53 μs

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1       239.61 KB - 4381.43x memory usage +239.55 KB

**All measurements for memory usage were the same**

##### With input 1200B valid #####
Name             ips        average  deviation         median         99th %
HEAD        671.39 K        1.49 μs    ±56.89%        1.42 μs        1.77 μs
HEAD^1       17.36 K       57.60 μs    ±33.50%       55.85 μs       72.75 μs

Comparison: 
HEAD        671.39 K
HEAD^1       17.36 K - 38.67x slower +56.11 μs

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1       239.61 KB - 4381.43x memory usage +239.55 KB

**All measurements for memory usage were the same**

Add fast-paths for fully valid and printable strings to return [string]
immediately, avoiding the chunk loop and allocations entirely for uniform
inputs.

In the fallback chunk loop:
- Reuse the first decoded codepoint and classification, eliminating
  redundant decoding and the intermediate wrapper.
- Replace repeated <> binary concatenation with zero-copy sub-binary
  slicing via binary_part/3, removing intermediate binary reallocations.

Assisted-by: Gemini:3.8-Flash
Assisted-by: Codex:GPT-6
@josevalim

Copy link
Copy Markdown
Member

Thanks!

This is an incomplete benchmark because very early in the string we see it is invalid:

{String.duplicate("hello" <> <<255, 254>> <> "world" <> <<128>>, 80), :valid},

We need to assert worst case scenarios like the invalid byte being at the end.

In any case, I am sure this could be much faster:

  • the valid case could traverse the binary using the same iteration as valid and either return nil (it is valid) or the rest. then you can get the byte_size of the rest, remove it from the original string, and start computing the next codepoints recursively. this means we don't need to traverse twice in worst cases like above

  • the printable case could have a similar optimization

Basically, I am saying we can likely remove the initial valid?/printable? check, anid that can be merged in the traversal only needs to compute the size of the next valid chunk.

@preciz

preciz commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Thank you for the feedback @josevalim, I investigate this more then.

@dkuku

dkuku commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

the valid case could traverse the binary using the same iteration as valid and either return nil (it is valid) or the rest.

@josevalim A bit off-topic, but I couldn't find an answer to this anywhere.

Could the type system keep track of whether a string is known to be valid UTF-8—for example, if it was already validated at compile time or during deserialization?

@josevalim

Copy link
Copy Markdown
Member

Yes, it could, but it is not a goal for now. For this pull request, I'd make it so chunking is close to a non-op until an invalid chunk is found, as proposed above.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants