Conversation
|
Did you try |
|
Next on my list of todo. I'll do it today and let you know. |
|
I think I fixed a lot in dev. Your feedback would be useful if I can promote these changes to master. |
|
Hi. Finally got around to testing. Claude told me the binary on my docker was about 18min old, so I'm not surprised you'd fixed some stuff. |
|
But I understood that problem existed only on serial connection. Local IP connection was always stable. |
|
I was definitely having many disconnections and timeouts every day using the cloud (sites id) connection. At some stage today I'll revert my PAI to dev without my changes and re-test and let you know. |
|
Here's the analysis based on my question: Q: Does the reference dev bench of PAI have the same adaptable timeout that we changed from default (0.5s) to 2.0 sec? We’re trying to work out if your changes to PAI have made any tangible improvements to connectivity stability issues I had. TLDR A: the changes improve cloud-based timeouts/errors, but less likely for direct IP connections. A: Only partly. I tested it directly: I ran upstream dev (without your changes) and your branch, each loading a config with IO_TIMEOUT = 2 the way pai-service starts up, and read back the timeouts each part of the code actually uses. What's being waited for Upstream dev (reference) Your branch (deployed) Upstream reads the setting for the IP150 part only by accident: that module happens to be loaded after the config. The waits for panel replies are fixed at 0.5 s when PAI starts, before your pai.conf is read. Your fix makes all of them read the setting at the time they wait. What that means for the problems you had: The honest caveat: yesterday I changed three things within two hours: Since the LAN switch 16¾ hours ago there have been no connection losses, no errors and no restarts. The warnings are down from about 2.6 a minute to 0.7 an hour. But on the LAN, replies take milliseconds, so the timeout barely matters any more, and the LAN change alone could account for most of the gain. The evidence so far proves your timeout fix matters over the cloud relay. It doesn't yet prove your changes make a difference on the setup you run now. Also, the old container only kept about 13 hours of logs, which show one dropout, so I don't have a solid "before" baseline for your original problem. |
Review of the connect/poll/reconnect paths for defects that cause or worsen repeated loss of connection to the panel. 13 findings, ranked by likely contribution, with file:line references against be1e46e. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…p stalls Addresses the 13 findings in CONNECTION_STABILITY_REVIEW.md. Tier 1: - IO_TIMEOUT was bound as a default argument, evaluated at import time -- before main() runs cfg.load() -- so every request path silently used the built-in 0.5 s and configuring it did nothing. Resolve it per call. - Bound the status poll cycle with a panel-owned budget and floor the idle gap between cycles, so an overrunning cycle no longer re-polls a struggling panel back to back with zero delay. - Reconnect backoff used '2 ^ retry' (XOR): 3, 0, 1, 6, 7, ... seconds, so the second attempt reconnected instantly and it never reached the cap. - Failed IP connect attempts left their socket open and unowned; the retry overwrote _protocol and PAI competed with its own orphans for the module's single session slot. Close each attempt, and pause between them. Tier 2 (STUN/paradoxmyhome): - The TURN refresh ran inline in write(), doing blocking socket I/O on the event loop with no socket timeout. Run it in an executor, bound the sockets, and drop the link when it fails. - Replaced a blocking time.sleep(5) in an async function, and bounded the SWAN site lookup. - receive_response() assumed one recv() returned a whole STUN message. Tier 3: - busy.release() ran in a finally that could not have acquired the lock. - gather left siblings running after the first status request failed. - Connection.close() skipped its state reset when the protocol raised -- which is the normal case when closing after a fault. - An unparsable status block took the whole cycle down with it. - disconnect() built a Connection just to ask whether one was open. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
I reverted back to your dev, left my timeout at 2s and switched back to cloud... What happened, step by step: What your PR changes here: On your branch, step 1 would still happen (the IP150 needs time), but the retries wouldn't trip over their own leftovers. Steps 2 and 3 shouldn't occur at all. |
a2cfe90 to
10febb4
Compare
|
Rebased onto current A/B run A: plain
So you were right: on
I didn't run B (this PR on the same relay setup). On 1 Oct I went back to a direct LAN connection to the IP150, where plain |
- S8572: log the STUN refresh failure with logger.exception() (same level
and traceback as the previous exc_info=True).
- S112: raise ConnectionError instead of Exception when the peer closes the
STUN control socket mid-response. The connect loop now reports it through
its OSError branch ("Connect failed") rather than as an unhandled exception.
- S5778: build the FutureHandler before pytest.raises, so only the awaited
call is inside it.
- S7483 (x3): mark the timeout parameter of wait_for_ip_message,
AsyncMessageManager.wait_for_message and Paradox.send_wait with NOSONAR.
It is the existing API (this PR only changes its default so IO_TIMEOUT is
read at call time), and the suggested asyncio.timeout() context manager
needs Python 3.11+ while PAI supports 3.8.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SonarCloud only accepts a bare "# NOSONAR" or "# NOSONAR(<rule keys>)"; the trailing explanation made the markers invalid (python:S7632) and left S7483 unsuppressed. Move the reason to its own comment line and scope each suppression to S7483. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|



What this is
My IP150/paradoxmyhome setup was dropping its panel connection several times a
day. Rather than patch around it locally I worked through the connection and
polling paths and ended up with 13 findings, most of which reproduce without any
hardware. This PR fixes them and adds regression tests for each.
Happy to split this into smaller PRs if that is easier to review — say the word
and I will break it up by tier.
Tier 1 — most likely to cause the drops
IO_TIMEOUTfrompai.confwas silently ignored.cfg.IO_TIMEOUTwascaptured into default arguments at import time, so the configured value never
reached any request path. Anyone who raised it to work around timeouts has
been running the default all along.
replies started arriving late. Now bounded by
Panel.status_cycle_budget; acycle that exceeds it is cancelled and counted as a missing reply, so it
surfaces as "Replies missing" and a reconnect instead of silence.
2 ^ retry— bitwise XOR, not exponentiation.The real sequence was 3, 0, 1, 6, 7… so the second retry fired immediately.
Now 2, 4, 8, 16, 30, 30 s.
IP150's single session slot so the retry could not get in.
Tier 2 — STUN / paradoxmyhome path
refresh_session_if_required()did blocking socket I/O on the event loop.time.sleep(5)inside an async function, plus an unbounded HTTP call.recv()returns a whole message.Tier 3 — smaller, still real
busy.release()could be called without holding the lock.asyncio.gatherabandoned sibling requests on first failure.disconnect()could construct a connection object during shutdown.Behaviour changes worth knowing about
outage, because PAI stops hammering the module.
CONNECT_RETRY_DELAY), so a failingconnect()takes ~10 s longer before handing back to the main retry loop.IO_TIMEOUTnow actually takes effect — existing configs shouldre-check their value.
SWAN site info instead of reusing a possibly stale
xoraddr. Costs one extraHTTPS round trip per retry.
into one request per area and zone.
Testing
21 new regression tests across
tests/connection/,tests/lib/,tests/paradox/andtests/test_main_uptime.py. Full suite for the touchedareas: 319 passed. Each finding has a test that fails before its fix.
Note on
CONNECTION_STABILITY_REVIEW.mdThe first commit adds the full write-up as a document in the repo root, with the
reasoning and file/line references behind each finding. I have kept it because it
makes the second commit reviewable, but I am happy to drop that commit if you
would rather it lived only in this PR description.