Skip to content

Fixes for v2.5.4: crash on fork+exec (with tests for script-running), printf specifier, pass correct stripc names - #638

Merged
paulusmack merged 4 commits into
ppp-project:masterfrom
keragez:master
Oct 1, 2026
Merged

paulusmack merged 4 commits into
ppp-project:masterfrom
keragez:master

Conversation

@keragez

@keragez keragez commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Situation

When running pppd v2.5.4 on an ppc64 system I encountered a crash of the pppd.

Core was generated by `/usr/sbin/pppd lcp-echo-failure 3 lcp-echo-interval 1 holdoff 0 mtu 1500 mru 1500 \
     nodefaultroute persist maxfail 0 lcp-echo-adaptive lcp-restart 3 lcp-max-configure 10 192.168.2.113: \
   ipv6cp-use-persistent noaccomp plugin /usr/lib/pppd/pppoe.so pppoe-sess 1:00:80:ea:aa:3c:fc unit 1 ppp'.
Program terminated with signal SIGSEGV, Segmentation fault.

Crash reproduces also on x86-64bit.

Crash after fork

The crash occured in the forked process in the pppd/main.c file around the ppp_script_setenv("PPP_SCRIPT_INSTANCE", name, 0); call. The execution never reached the execve(...), whereby the forked pppd instance was never superseeded by the script it intended to run.

Analysis of core suggested that the pppdb was closed two times. Although int tdb_close(TDB_CONTEXT *tdb) uses SAFE_FREE for the actual freeing, it unconditionally dereferences tdb->map_ptr, so the pppdb ptr must be NULLed after freeing.

Added test for the fix, to prevent further regressions.

The tests:

  • skip on non-TDB builds
  • fail on TDB builds without the fix for pppdb NULL
  • pass on TDB builds with the fix for pppdb NULL

Non-TDB build:

~/work/ppp$ grep --text --count pppd2.tdb pppd/pppd
0
~/work/ppp $ sudo ./runtests.py script-run
============================================================
./runtests.py running in /home/adamz/work/ppp
    pppd_bin=/home/adamz/work/ppp/pppd/pppd
    srcdir=/home/adamz/work/ppp
    os=Linux gdn-s-csl8 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug  1 03:52:05 UTC 2026 x86_64 Intel(R) Xeon(R) CPU E5-1660 0 @ 3.30GHz GenuineIntel GNU/Linux
    preserve_scratch=no
    scratchbase=/home/adamz/work/ppp/testtmp
SKIP    script-run (/home/adamz/work/ppp/pppd/pppd built without TDB support (--enable-multilink))
------------------------------------------------------------
----- overall results:
      0 passed
      1 skipped
------------------------------------------------------------
overall result is 0

TDB build, no pppdb fix:

~/work/ppp$ grep --text --count pppd2.tdb pppd/pppd
1
~/work/ppp$ sudo ./runtests.py script-run
============================================================
./runtests.py running in /home/adamz/work/ppp
    pppd_bin=/home/adamz/work/ppp/pppd/pppd
    srcdir=/home/adamz/work/ppp
    os=Linux gdn-s-csl8 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug  1 03:52:05 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux
    preserve_scratch=no
    scratchbase=/home/adamz/work/ppp/testtmp
----- script-run log follows
/etc/ppp/auth-down did not run within 30s
pppd reported a dying script child: Child process /usr/local/etc/ppp/auth-down (pid 2221847) terminated with signal 11
ip-pre-up / ip-up / ip-down:
hello from ip-pre-up instance=ip-pre-up args=ppp0  38400 10.0.0.1 10.0.0.2
hello from ip-up instance=ip-up args=ppp0  38400 10.0.0.1 10.0.0.2
hello from ip-down instance=ip-down args=ppp0  38400 10.0.0.1 10.0.0.2
auth-up / auth-down:
hello from auth-up instance=auth-up args=ppp0 cli root  38400
----- script-run log ends
----- script-run auth-a/pppd.log follows
using channel 1
Using interface ppp0
Connect: ppp0 <--> /dev/pts/36
sent [LCP ConfReq id=0x1 <asyncmap 0x0> <auth pap> <magic 0x2bae40ba> <pcomp> <accomp>]
rcvd [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0x6a88001c> <pcomp> <accomp>]
sent [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0x6a88001c> <pcomp> <accomp>]
rcvd [LCP ConfAck id=0x1 <asyncmap 0x0> <auth pap> <magic 0x2bae40ba> <pcomp> <accomp>]
sent [LCP EchoReq id=0x0 magic=0x2bae40ba]
rcvd [LCP EchoReq id=0x0 magic=0x6a88001c]
sent [LCP EchoRep id=0x0 magic=0x2bae40ba]
rcvd [PAP AuthReq id=0x1 user="cli" password=<hidden>]
sent [PAP AuthAck id=0x1 "Login ok"]
PAP peer authentication succeeded for cli
Peer cli authenticated with PAP
Script /usr/local/etc/ppp/auth-up started (pid 2221845)
sent [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
sent [IPV6CP ConfReq id=0x1 <addr fe80::5917:b463:3dab:f730>]
rcvd [LCP EchoRep id=0x0 magic=0x6a88001c]
rcvd [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
rcvd [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
sent [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
rcvd [IPV6CP ConfReq id=0x1 <addr fe80::11a0:b432:6ffe:8026>]
sent [IPV6CP ConfAck id=0x1 <addr fe80::11a0:b432:6ffe:8026>]
rcvd [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
Deflate (15) compression enabled
rcvd [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
local  IP address 10.0.0.1
remote IP address 10.0.0.2
Script /usr/local/etc/ppp/auth-up finished (pid 2221845), status = 0x0
rcvd [IPV6CP ConfAck id=0x1 <addr fe80::5917:b463:3dab:f730>]
local  LL address fe80::5917:b463:3dab:f730
remote LL address fe80::11a0:b432:6ffe:8026
rcvd [LCP TermReq id=0x2 "User request"]
LCP terminated by peer (User request)
Script /usr/local/etc/ppp/auth-down started (pid 2221847)
Connect time 0.0 minutes.
Sent 48 bytes, received 48 bytes.
sent [LCP TermAck id=0x2]
Modem hangup
Connection terminated.
Script pppd (charshunt) finished (pid 2221835), status = 0x0
Waiting for 1 child processes...
  script /usr/local/etc/ppp/auth-down, pid 2221847
Child process /usr/local/etc/ppp/auth-down (pid 2221847) terminated with signal 11
----- script-run auth-a/pppd.log ends
----- script-run auth-b/pppd.log follows
using channel 1
Using interface ppp0
Connect: ppp0 <--> /dev/pts/38
sent [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0x6a88001c> <pcomp> <accomp>]
rcvd [LCP ConfReq id=0x1 <asyncmap 0x0> <auth pap> <magic 0x2bae40ba> <pcomp> <accomp>]
sent [LCP ConfAck id=0x1 <asyncmap 0x0> <auth pap> <magic 0x2bae40ba> <pcomp> <accomp>]
rcvd [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0x6a88001c> <pcomp> <accomp>]
sent [LCP EchoReq id=0x0 magic=0x6a88001c]
sent [PAP AuthReq id=0x1 user="cli" password=<hidden>]
rcvd [LCP EchoReq id=0x0 magic=0x2bae40ba]
sent [LCP EchoRep id=0x0 magic=0x6a88001c]
rcvd [LCP EchoRep id=0x0 magic=0x2bae40ba]
rcvd [PAP AuthAck id=0x1 "Login ok"]
Remote message: Login ok
PAP authentication succeeded
sent [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
sent [IPV6CP ConfReq id=0x1 <addr fe80::11a0:b432:6ffe:8026>]
rcvd [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
rcvd [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
sent [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
rcvd [IPV6CP ConfReq id=0x1 <addr fe80::5917:b463:3dab:f730>]
sent [IPV6CP ConfAck id=0x1 <addr fe80::5917:b463:3dab:f730>]
rcvd [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
Deflate (15) compression enabled
rcvd [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
local  IP address 10.0.0.2
remote IP address 10.0.0.1
rcvd [IPV6CP ConfAck id=0x1 <addr fe80::11a0:b432:6ffe:8026>]
local  LL address fe80::11a0:b432:6ffe:8026
remote LL address fe80::5917:b463:3dab:f730
Terminating on signal 15
Connect time 0.0 minutes.
Sent 48 bytes, received 48 bytes.
sent [LCP TermReq id=0x2 "User request"]
rcvd [LCP TermAck id=0x2]
Connection terminated.
Waiting for 1 child processes...
  script pppd (charshunt), pid 2221844
Script pppd (charshunt) finished (pid 2221844), status = 0x0
----- script-run auth-b/pppd.log ends
----- script-run ip-a/pppd.log follows
using channel 1
Using interface ppp0
Connect: ppp0 <--> /dev/pts/38
sent [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0xd17fed1d> <pcomp> <accomp>]
rcvd [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0xb297d762> <pcomp> <accomp>]
sent [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0xb297d762> <pcomp> <accomp>]
rcvd [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0xd17fed1d> <pcomp> <accomp>]
sent [LCP EchoReq id=0x0 magic=0xd17fed1d]
sent [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
sent [IPV6CP ConfReq id=0x1 <addr fe80::806f:1f24:d83d:d3a5>]
rcvd [LCP EchoReq id=0x0 magic=0xb297d762]
sent [LCP EchoRep id=0x0 magic=0xd17fed1d]
rcvd [LCP EchoRep id=0x0 magic=0xb297d762]
rcvd [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
rcvd [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
sent [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
rcvd [IPV6CP ConfReq id=0x1 <addr fe80::e12c:6445:2506:a945>]
sent [IPV6CP ConfAck id=0x1 <addr fe80::e12c:6445:2506:a945>]
rcvd [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
Deflate (15) compression enabled
rcvd [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
Script /usr/local/etc/ppp/ip-pre-up started (pid 2221814)
Script /usr/local/etc/ppp/ip-pre-up finished (pid 2221814), status = 0x0
local  IP address 10.0.0.1
remote IP address 10.0.0.2
Script /usr/local/etc/ppp/ip-up started (pid 2221815)
rcvd [IPV6CP ConfAck id=0x1 <addr fe80::806f:1f24:d83d:d3a5>]
local  LL address fe80::806f:1f24:d83d:d3a5
remote LL address fe80::e12c:6445:2506:a945
Script /usr/local/etc/ppp/ip-up finished (pid 2221815), status = 0x0
rcvd [LCP TermReq id=0x2 "User request"]
LCP terminated by peer (User request)
Connect time 0.0 minutes.
Sent 48 bytes, received 0 bytes.
Script /usr/local/etc/ppp/ip-down started (pid 2221820)
sent [LCP TermAck id=0x2]
Script /usr/local/etc/ppp/ip-down finished (pid 2221820), status = 0x0
Modem hangup
Connection terminated.
Script pppd (charshunt) finished (pid 2221807), status = 0x0
----- script-run ip-a/pppd.log ends
----- script-run ip-b/pppd.log follows
using channel 1
Using interface ppp0
Connect: ppp0 <--> /dev/pts/36
sent [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0xb297d762> <pcomp> <accomp>]
rcvd [LCP ConfReq id=0x1 <asyncmap 0x0> <magic 0xd17fed1d> <pcomp> <accomp>]
sent [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0xd17fed1d> <pcomp> <accomp>]
rcvd [LCP ConfAck id=0x1 <asyncmap 0x0> <magic 0xb297d762> <pcomp> <accomp>]
sent [LCP EchoReq id=0x0 magic=0xb297d762]
sent [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
sent [IPV6CP ConfReq id=0x1 <addr fe80::e12c:6445:2506:a945>]
rcvd [LCP EchoReq id=0x0 magic=0xd17fed1d]
sent [LCP EchoRep id=0x0 magic=0xb297d762]
rcvd [CCP ConfReq id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
sent [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
rcvd [LCP EchoRep id=0x0 magic=0xd17fed1d]
rcvd [IPCP ConfReq id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
sent [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.1>]
rcvd [IPV6CP ConfReq id=0x1 <addr fe80::806f:1f24:d83d:d3a5>]
sent [IPV6CP ConfAck id=0x1 <addr fe80::806f:1f24:d83d:d3a5>]
rcvd [CCP ConfAck id=0x1 <deflate 15> <deflate(old#) 15> <bsd v1 15>]
Deflate (15) compression enabled
rcvd [IPCP ConfAck id=0x1 <compress VJ 0f 01> <addr 10.0.0.2>]
local  IP address 10.0.0.2
remote IP address 10.0.0.1
rcvd [IPV6CP ConfAck id=0x1 <addr fe80::e12c:6445:2506:a945>]
local  LL address fe80::e12c:6445:2506:a945
remote LL address fe80::806f:1f24:d83d:d3a5
Terminating on signal 15
Connect time 0.0 minutes.
Sent 48 bytes, received 48 bytes.
sent [LCP TermReq id=0x2 "User request"]
rcvd [LCP TermAck id=0x2]
Connection terminated.
Waiting for 1 child processes...
  script pppd (charshunt), pid 2221806
Script pppd (charshunt) finished (pid 2221806), status = 0x0
----- script-run ip-b/pppd.log ends
FAIL    script-run
------------------------------------------------------------
----- overall results:
      0 passed
      1 failed
------------------------------------------------------------
overall result is 1

TDB build, with pppdb fix:

~/work/ppp$ grep --text --count pppd2.tdb pppd/pppd
1
~/work/ppp$ sudo ./runtests.py script-run
============================================================
./runtests.py running in /home/adamz/work/ppp
    pppd_bin=/home/adamz/work/ppp/pppd/pppd
    srcdir=/home/adamz/work/ppp
    os=Linux gdn-s-csl8 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug  1 03:52:05 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux
    preserve_scratch=no
    scratchbase=/home/adamz/work/ppp/testtmp
PASS    script-run
------------------------------------------------------------
----- overall results:
      1 passed
------------------------------------------------------------
overall result is 0

Other issues

Analisys of differences from v2.5.2 also shown:

  • wrong format specifier on .*B in a dbglog function call.
  • wrong script names used when calling auth-down, ipv6-up and ip-up

-BR,
AZ

@keragez

keragez commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Hello, this PR is similar to PR#637 in that that it fixes the pppdb NULL pointer issue after closing the pppdb, but it also:

  • adds test to guard against regression,
  • addresses format string error,
  • fixes names in script calling.

The debug message in eaptls_send() uses "%d ... %.*B", which consumes
three arguments: the byte count for %d, then an int precision and a
buffer pointer for %.*B. Only two were passed, so the count was taken
as the %d, the dummy buffer pointer was taken as the precision, and
vslprintf() fetched the %B data pointer from whatever followed on the
argument list. With "debug" enabled this reads through an undefined
pointer for an effectively arbitrary number of bytes and can crash
pppd during an EAP-TLS handshake.

Pass res for %d and MIN(res, 20) as the precision, as intended.

Fixes: dd5acd9 ("pppd/EAP-TLS: Send zero byte as protected success indication with TLS 1.3")
Signed-off-by: Adam Zegarek <keragez@gmail.com>
…up and auth-down

run_program() exports its name argument to the script as
PPP_SCRIPT_INSTANCE. pppd.8 documents it as "the name of the intended
script, as documented, not as referenced", and recommends it to scripts
because with strict-script-checks (the default) a #! script is run via
fexecve() and its $0 no longer identifies which hook it is.

Three call sites pass a name that does not match the script they run:

 - link_down() runs path_auth_down labelled "auth-up". This is the
   normal link-down path whenever the peer authenticated.
 - ipv6cp_up() runs path_ipv6up labelled "ipv6-ip". This is every
   time IPv6CP comes up.
 - ipcp_script_done() runs path_ipup labelled "ip-down" when the link
   came back up while ip-down was still running.

The correct script file was always executed; only the environment
variable was wrong. A single script installed for several hooks that
dispatches on $PPP_SCRIPT_INSTANCE would take the wrong branch, e.g.
re-adding rules on auth-down, or doing nothing on ipv6-up.

Fixes: e87ddf0 ("run_program:  Export the name of the script via PPP_SCRIPT_INSTANCE")
Signed-off-by: Adam Zegarek <keragez@gmail.com>
The forked child closes the TDB with tdb_close(), which zeroes and
frees the context, but leaves the global pppdb pointing at the freed
memory.

This was harmless until e87ddf0, because nothing in the child touched
pppdb before execve(). run_program() now calls
ppp_script_setenv("PPP_SCRIPT_INSTANCE", ...) in the child, which sees
pppdb != NULL and calls update_db_entry() -> tdb_store() on the freed
context. The allocation of the new environment string can reuse that
chunk, so the result depends on heap layout: often the child survives,
sometimes it dies with SIGSEGV before it executes the script. The
parent only logs "Child process ... terminated with signal 11" and
carries on, so ip-up, ip-down, auth-up, auth-down etc. are silently
skipped.

Seen in the field on 2.5.4 with PPPoE as SIGSEGV core dumps of pppd
(the script child), and reproduced on x86_64 with the new script-run
test, where the auth-down child crashed.

Only builds with --enable-multilink (PPP_WITH_TDB) are affected.

Fixes: e87ddf0 ("run_program:  Export the name of the script via PPP_SCRIPT_INSTANCE")
Signed-off-by: Adam Zegarek <keragez@gmail.com>

@jkroonza jkroonza left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably missing something obvious, but please confirm if there is a reason to skip the run_program() tests with TDB disabled - surely the scripts should work with it enabled as well as disabled?

I know the regression was only with it enabled, but we can test both ways right?

Otherwise looks all good for me.

@keragez

keragez commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

confirm if there is a reason to skip the run_program() tests with TDB disabled

The reason was: to have a test in place to look for regression regrading the pppdb NULL - if it impacts run_program(). I didn't want to test run_program() itself, I wanted a to look out for regression in pppdb.

So, if you haven't build ppp-project with PPP_WITH_TDB enabled, there is not pppdb regression to look out for.

I know the regression was only with it enabled, but we can test both ways right?

Well... You're right. It would still check script-running and having correct names in script-calling. Changed skipping to execution.

Non-TDB build:

adamz@gdn-s-csl8 ~/work/ppp$ grep --text --count pppd2.tdb pppd/pppd
0
adamz@gdn-s-csl8 ~/work/ppp$ sudo ./runtests.py script-run
============================================================
./runtests.py running in /home/adamz/work/ppp_keragez
    pppd_bin=/home/adamz/work/ppp_keragez/pppd/pppd
    srcdir=/home/adamz/work/ppp_keragez
    os=Linux gdn-s-csl8 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug  1 03:52:05 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux
    preserve_scratch=no
    scratchbase=/home/adamz/work/ppp_keragez/testtmp
PASS    script-run
------------------------------------------------------------
----- overall results:
      1 passed
------------------------------------------------------------
overall result is 0
~/work/ppp$ sudo ./runtests.py script-run --always-log | grep TDB
note: /home/adamz/work/ppp/pppd/pppd built without TDB (--enable-multilink); pppdb regression not exercised

-AZ

@keragez

keragez commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

@jkroonza: So, we're going forward with this patch?

@jkroonza

Copy link
Copy Markdown
Contributor

@keragez yes. I'm not sure what the time-lines are that @paulusmack is aiming for a potential follow-up release, I'll include on Gentoo side so long as this is going to potentially bite hard. Would it be useful to introduce a tdb USE flag in your opinion?

@keragez

keragez commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

Oh, did the ci# test discovered yet another regression?
On OmniOS pppd cannot execute scripts?

The tests we just enabled helped discover it. I'll look into that.

Would it be useful to introduce a tdb USE flag in your opinion?

What's the benefit in that?
Does it limit the size of the binary by a lot?

TDB is enabled if and only if --enable-multilink is. So, shoulnd't the use flag be name multilink, to follow the feature it controls? It would have to be default enabled, not to break existing setups. And turning tdb to off does not fix the bug, jut hides it.

-AZ

@jkroonza

Copy link
Copy Markdown
Contributor

2.5.4 with TDB there are cases where the scripts won't execute. This is added to the test suite now. The cause was the TDB dangling pointer in the child, which was referenced - if this pointer referenced invalid data we potentially just got the lock error, if it was corrupt in a different way we could crash (and then the scripts would not execute).

Let me get the patch into gentoo as -r1. Then I'll look into the USE flag (and yes, I think multilink is more appropriate). Either way - looks like ppp uses a bundled, potentially outdated tdb. @paulusmack would it be possible to enable using the system-installed tdb?

@keragez

keragez commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

As to the OmniOS CI failure

The OmniOS failure is not caused by this PR. The new test surfaced an existing 2.5.4 problem that no test had covered before.

What happened

After the change to also run script-run without TDB, the test ran on OmniOS for the first time (multilink/TDB isn't available on SunOS, so it used to skip there). pppd forks every hook, but every one exits with status 0x63 on both peers:

Script /usr/local/etc/ppp/ip-pre-up started (pid 3818)
Script /usr/local/etc/ppp/ip-pre-up finished (pid 3818), status = 0x63

0x63 = 99 is the _exit(99) in run_program() after the exec call fails. So the script never starts, and the child doesn't crash (no "terminated with signal").

Likely cause

Since 6a4944f ("pppd: relax and simplify permission check", first released in 2.5.4), run_program() opens the script with O_PATH, or O_EXEC where there's no O_PATH, and under strict-script-checks (the default) runs it with fexecve(fd). On Linux (O_PATH) that works for #! scripts. On illumos, configure finds fexecve, the descriptor is O_EXEC, and the exec fails. 2.5.2 and 2.5.3 exec'd by path.

I don't have the exact errno: pppd logs Can't execute …: %m to syslog, and the CI job doesn't print syslog.

Solaris 11.4 passes only because it has no kernel PPP, so all link tests skip there. Linux CI is unaffected.

I also don't have an OmniOS machine to fix and test it on locally.

What I changed here

  • script-run still runs on every build (with and without TDB) as suggested. On SunOS, the specific "hook not executed, pppd logged exit status 0x63" case is reported as XFAIL. Any other failure, including a script child dying on a signal (the pppdb crash this PR fixes), still fails. If hooks start working on illumos, the test simply passes and the XFAIL branch can be removed.
  • The failure message now shows the real confdir path instead of a hardcoded /etc/ppp/.
  • Corrected my earlier commit message: illumos uses fexecve(), not the /dev/fd/N fallback.

@jkroonza : This would be a separate issue for the illumos exec problem. Fixing it means choosing how strict-script-checks should behave without O_PATH, and I'd rather not fold that into this PR. Should I raise it?

@jkroonza

jkroonza commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Please raise a separate PR/issue for that. You can point to this PR for the tests. @paulusmack may I request this get merged sooner rather than later please. I think a 2.5.5 may be in order once this, #635 and a fix for the OmniOS issues are merged. I'll see if I can bring an OmniOS up somewhere on a VM. Need to bring something up for another purpose anyway.

I am going to recommend forking a 2.5 branch at this point with for further fixes pertaining to 2.5.X, with any new features only going to master, targeting that for 2.6.0 eventually some time 2027q1 probably. Things like radius-ng will then also be targeted at that rather than bringing into 2.5.X which I then highly recommend gets a "security & bug fix only" status associated.

I think for a 2.6 we should try and focus strongly on increasing test coverage as much as possible.

gentoo-bot pushed a commit to gentoo/gentoo that referenced this pull request Sep 30, 2026
Include fairly critical fixes: ppp-project/ppp#638

Signed-off-by: Jaco Kroon <jkroon@gentoo.org>
@jkroonza

Copy link
Copy Markdown
Contributor

Two things @keragez

  1. Out of interest - how much AI was involved here? I don't think Paul has a policy against it, and frankly, I don't mind AI assisted, but I personally distrust anything purely done by AI where human was not overseeing and managing the process. Specifically the testsuite stuff looks to have some involvement. Sorry, I've seen way too many things go completely sideways thanks to AI. The main code fixes themselves looks perfectly fine.
  2. Could you please squash the last two commits? Treat the latter as a fixup for the first.

No existing test installs a hook script, and run_program() returns
before forking when the script file does not exist, so the code between
fork() and execve() was never run by the suite. That is how a SIGSEGV
in the script child (see "pppd: Clear pppdb after tdb_close() in
ppp_safe_fork()") went unnoticed.

pppfns.py: PppPeer gains a scripts={name: text} argument. The scripts
are staged exactly like pap-secrets/chap-secrets: on Linux they are
copied into the root-owned tmpfs behind the bind-mounted confdir and
symlinked from it, elsewhere they are copied into the real confdir
(PPPD_TEST_GLOBAL_CONF=1 hosts only) and removed afterwards. They are
installed mode 755 so ppp_check_access(PPP_FT_EXEC) accepts them.

script-run_test.py installs hooks that append their name,
$PPP_SCRIPT_INSTANCE and arguments to a marker file, then:

 - brings a link up and down and checks ip-pre-up (run with wait=1),
   ip-up and ip-down (wait=0) all ran;
 - repeats with side a as a PAP server and checks auth-up and
   auth-down ran;
 - checks PPP_SCRIPT_INSTANCE matches each hook (skipped when unset,
   so --pppd-bin2 against pre-2.5.4 pppd still works);
 - fails if pppd logged a script child terminated by a signal.

The pppdb crash can only happen in builds with TDB
(--enable-multilink), but the test runs on every build: the hooks must
run and be labelled correctly regardless, and the PPP_SCRIPT_INSTANCE
mislabelling fixed in "pppd: Pass correct PPP_SCRIPT_INSTANCE name for
deferred ip-up, ipv6-up and auth-down" affected every build. Without
TDB it prints a note, so that a pass there is not mistaken for coverage
of the pppdb fix. Running everywhere also covers Solaris/illumos, where
multilink is not supported: illumos has no O_PATH, so run_program()
fexecve()s a descriptor opened with O_EXEC, a path the Linux CI jobs
never take.

On OmniOS every hook currently exits with status 0x63 (99, the child's
exit after a failed exec): since 2.5.4, under the default
strict-script-checks, run_program() fails to fexecve() #! scripts from
an O_EXEC descriptor there, while 2.5.3 and earlier exec'd by path.
That is a separate issue, not what this test guards, so on SunOS a hook
that did not run *and* was logged exiting 0x63 is reported as XFAIL.
Every other failure, including a script child killed by a signal, still
fails; if hooks start working on illumos the test just passes.

The test is skipped on non-Linux hosts that already have one of the
hook scripts installed. The marker file is created by the test user
beforehand, since the script child runs with umask 077.

Results on x86_64 Linux:
  without TDB:                PASS (with note)
  TDB, without pppdb fix:     FAIL (auth-down child terminated with signal 11)
  TDB, with pppdb fix:        PASS

Signed-off-by: Adam Zegarek <keragez@gmail.com>
@keragez

keragez commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Ad.1: I have Opus 5.5 and used it. I employed it mainly for the things these machines are good for - for reading looooong inputs and summarizing.

We had a regression from v2.5.2 to v2.5.4, so I looked up the diff myself, and then fed it the diff to look for potential issues. It had failed to find TDB issue first, but found the .*B format issue right off the bat. So, I got the core and fed the backtrace to the machine, and we landed together on the right issue.

Tests - mainly AI driven, but read through and tested by me before submitting. I also fed it your Submitting-patches.md policy, and it added some verbiage to commit messages, to make them pass the policy of being "self sufficient without Github's infrastructure".

Ad. 2: done, squash pushed, all tests in one commit.

@paulusmack

Copy link
Copy Markdown
Collaborator

Thanks @keragez, this is great. :)

@paulusmack
paulusmack merged commit 13ba8a5 into ppp-project:master Oct 1, 2026
33 checks passed
@paulusmack

Copy link
Copy Markdown
Collaborator

Let me get the patch into gentoo as -r1. Then I'll look into the USE flag (and yes, I think multilink is more appropriate). Either way - looks like ppp uses a bundled, potentially outdated tdb. @paulusmack would it be possible to enable using the system-installed tdb?

That's a good idea. Alternatively we could change to sqlite or similar.

@jkroonza

jkroonza commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Let me get the patch into gentoo as -r1. Then I'll look into the USE flag (and yes, I think multilink is more appropriate). Either way - looks like ppp uses a bundled, potentially outdated tdb. @paulusmack would it be possible to enable using the system-installed tdb?

That's a good idea. Alternatively we could change to sqlite or similar.

Let's not create too many options. I think we should also keep backwards compatibility in mind. Consider the use case where I upgrade ppp but don't want to reconnect all dial-in users just to upgrade the database. In this case it's better to just stick everything with TDB. Let's log a separate issue for this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants