Skip to content

Add FASTQ header cleaning as the first pipeline step - #35

Open
Chromojones wants to merge 1 commit into
goodwright:masterfrom
Chromojones:clean-fastq-headers
Open

Chromojones wants to merge 1 commit into
goodwright:masterfrom
Chromojones:clean-fastq-headers

Conversation

@Chromojones

@Chromojones Chromojones commented Sep 8, 2026 •

Copy link
Copy Markdown

Some FASTQ files carry spaces in their header lines (for example "ID 1:N:0:INDEX"). Tools further down the pipeline truncate read names at the first space, which breaks UMI extraction and read-name matching after alignment (samtools).

CLEAN_FASTQ_HEADERS replaces every space in a header line with an underscore. It runs on the raw reads immediately after the samplesheet is parsed, before UMI moving, trimming and alignment, and reassigns ch_fastq the same way the surrounding steps do. Sequence and quality lines are passed through untouched.

Off by default, enable with --clean_fastq_headers.

The cleaned reads are intermediates and are not published. TrimGalore re-links its input to "${prefix}.fastq.gz", so downstream published file names are unchanged. pigz signals a corrupt or truncated input by exiting non-zero after emitting a partial stream, so both ends of the pipe are checked and the record count is validated.

Some FASTQ files carry spaces in their header lines (for example
"@Readid 1:N:0:INDEX"). Tools further down the pipeline truncate read
names at the first space, which breaks UMI extraction and read-name
matching after alignment.

CLEAN_FASTQ_HEADERS replaces every space in a header line with an
underscore. It runs on the raw reads immediately after the samplesheet
is parsed, before UMI moving, trimming and alignment, and reassigns
ch_fastq the same way the surrounding steps do, so nothing downstream
needed changing. Sequence and quality lines are passed through
untouched.

Off by default; enable with --clean_fastq_headers.

The cleaned reads are intermediates and are not published. TrimGalore
re-links its input to "${prefix}.fastq.gz", so downstream published
file names are unchanged.

pigz signals a corrupt or truncated input by exiting non-zero after
emitting a partial stream, so both ends of the pipe are checked and the
record count is validated; otherwise the task would succeed and hand a
silently truncated FASTQ to the rest of the pipeline.

The pipeline is single-end only - single_end is parsed into meta but
never read downstream - so the module accepts exactly one FASTQ per
sample and fails loudly on anything else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant