Skip to content

Optimize String.chunk with uniform fast-path and sub-binary slicing - #15958

Closed
preciz wants to merge 1 commit into
elixir-lang:mainfrom
preciz:perf/string-chunk-reuse
Closed

preciz wants to merge 1 commit into
elixir-lang:mainfrom
preciz:perf/string-chunk-reuse

Conversation

@preciz

@preciz preciz commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Assisted-by: Antigravity CLI:Gemini 3.8 Flash

Add fast-paths for 100% valid and printable strings to return [string]
immediately, avoiding the chunk loop in those cases.

In the fallback chunk loop:

  • Reuse the first decoded codepoint and classification, eliminating
    redundant decoding and the intermediate wrapper.
  • Replace repeated <> binary concatenation with zero-copy sub-binary
    slicing via binary_part/3, removing intermediate binary reallocations.

Bench:

# Run with: elixir bench.exs
Mix.install([:benchee])

path = "lib/elixir/lib/string.ex"

for {ref, module} <- [{"HEAD^1", Old}, {"HEAD", New}] do
  {source, 0} = System.cmd("git", ["show", "#{ref}:#{path}"], cd: __DIR__)

  source
  |> String.replace("defmodule String do", "defmodule #{inspect(module)} do")
  |> Code.compile_string(path)
end

Benchee.run(
  %{
    "HEAD^1" => fn {string, trait} -> Old.chunk(string, trait) end,
    "HEAD" => fn {string, trait} -> New.chunk(string, trait) end
  },
  inputs: %{
    "11B valid" => {"hello world", :valid},
    "11B printable" => {"hello world", :printable},
    "1200B valid" =>
      {String.duplicate("The quick brown fox jumps over the lazy dog! ✨ ", 25), :valid},
    "1200B printable" =>
      {String.duplicate("The quick brown fox jumps over the lazy dog! ✨ ", 25), :printable},
    "1200B mixed valid/invalid" =>
      {String.duplicate("hello" <> <<255, 254>> <> "world" <> <<128>>, 80), :valid},
    "1200B mixed printable" =>
      {String.duplicate("abc\x00def\tghi\x1b[0m", 80), :printable}
  },
  pre_check: :all_same,
  warmup: 1,
  time: 3,
  memory_time: 1
)

Results:

Operating System: Linux
CPU Information: AMD Ryzen 7 8845HS w
Number of Available Cores: 16
Available memory: 54.72 GB
Elixir 1.20.4
Erlang 29.0.5
JIT enabled: true

Benchmark suite executing with the following configuration:
warmup: 1 s
time: 3 s
memory time: 1 s
reduction time: 0 ns
parallel: 1
inputs: 11B printable, 11B valid, 1200B mixed printable, 1200B mixed valid/invalid, 1200B printable, 1200B valid
Estimated total run time: 1 min
Excluding outliers: false

##### With input 11B printable #####
Name             ips        average  deviation         median         99th %
HEAD         13.42 M       74.50 ns   ±474.77%          70 ns         170 ns
HEAD^1        1.17 M      856.65 ns   ±721.97%         771 ns        1543 ns

Comparison: 
HEAD         13.42 M
HEAD^1        1.17 M - 11.50x slower +782.14 ns

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1         2.30 KB - 42.14x memory usage +2.25 KB

**All measurements for memory usage were the same**

##### With input 11B valid #####
Name             ips        average  deviation         median         99th %
HEAD         16.83 M       59.41 ns  ±2402.20%          50 ns          90 ns
HEAD^1        1.34 M      748.97 ns   ±901.97%         641 ns        1522 ns

Comparison: 
HEAD         16.83 M
HEAD^1        1.34 M - 12.61x slower +689.56 ns

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1         2.30 KB - 42.14x memory usage +2.25 KB

**All measurements for memory usage were the same**

##### With input 1200B mixed printable #####
Name             ips        average  deviation         median         99th %
HEAD         22.21 K       45.02 μs    ±31.57%       43.91 μs       52.02 μs
HEAD^1       11.65 K       85.80 μs    ±35.88%       80.39 μs      144.72 μs

Comparison: 
HEAD         22.21 K
HEAD^1       11.65 K - 1.91x slower +40.78 μs

Memory usage statistics:

Name      Memory usage
HEAD         202.66 KB
HEAD^1       245.34 KB - 1.21x memory usage +42.69 KB

**All measurements for memory usage were the same**

##### With input 1200B mixed valid/invalid #####
Name             ips        average  deviation         median         99th %
HEAD         24.51 K       40.80 μs    ±35.33%       39.56 μs       50.85 μs
HEAD^1       14.48 K       69.06 μs    ±43.98%       59.84 μs      167.18 μs

Comparison: 
HEAD         24.51 K
HEAD^1       14.48 K - 1.69x slower +28.26 μs

Memory usage statistics:

Name      Memory usage
HEAD         181.38 KB
HEAD^1       215.33 KB - 1.19x memory usage +33.95 KB

**All measurements for memory usage were the same**

##### With input 1200B printable #####
Name             ips        average  deviation         median         99th %
HEAD        228.93 K        4.37 μs    ±45.77%        4.28 μs        5.56 μs
HEAD^1       13.90 K       71.95 μs    ±32.30%       70.42 μs       79.90 μs

Comparison: 
HEAD        228.93 K
HEAD^1       13.90 K - 16.47x slower +67.58 μs

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1       239.61 KB - 4381.43x memory usage +239.55 KB

**All measurements for memory usage were the same**

##### With input 1200B valid #####
Name             ips        average  deviation         median         99th %
HEAD        694.59 K        1.44 μs    ±63.08%        1.37 μs        1.73 μs
HEAD^1       19.23 K       52.00 μs    ±35.20%       50.45 μs       60.07 μs

Comparison: 
HEAD        694.59 K
HEAD^1       19.23 K - 36.12x slower +50.56 μs

Memory usage statistics:

Name      Memory usage
HEAD         0.0547 KB
HEAD^1       239.61 KB - 4381.43x memory usage +239.55 KB

**All measurements for memory usage were the same**

Add fast-paths for 100% valid and printable strings to return [string]
immediately, avoiding the chunk loop and allocations entirely for uniform
inputs.

In the fallback chunk loop:
- Reuse the first decoded codepoint and classification, eliminating
  redundant decoding and the intermediate wrapper.
- Replace repeated <> binary concatenation with zero-copy sub-binary
  slicing via binary_part/3, removing intermediate binary reallocations.

Benchee / OTP 29 (1s warmup, 3s time, 1s memory, pre_check: :all_same):

| Input | Time parent -> patch (speedup) | Memory parent -> patch (change) |
| --- | ---: | ---: |
| 11B valid | 748.97 ns -> 59.41 ns (12.61x) | 2.30 KB -> 0.05 KB (-97.6%) |
| 11B printable | 856.65 ns -> 74.50 ns (11.50x) | 2.30 KB -> 0.05 KB (-97.6%) |
| 1200B valid | 52.00 μs -> 1.44 μs (36.12x) | 239.61 KB -> 0.05 KB (-99.98%) |
| 1200B printable | 71.95 μs -> 4.37 μs (16.47x) | 239.61 KB -> 0.05 KB (-99.98%) |
| 1200B mixed valid/invalid | 69.06 μs -> 40.80 μs (1.69x) | 215.33 KB -> 181.38 KB (-15.8%) |
| 1200B mixed printable | 85.80 μs -> 45.02 μs (1.91x) | 245.34 KB -> 202.66 KB (-17.4%) |

Assisted-by: Gemini:3.8-Flash
@preciz

preciz commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Needs more work, closing it for now.

@preciz preciz closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant