160fff665e326952ed35aea9be5ed93cefae9c2e - haproxy

commit	160fff665e326952ed35aea9be5ed93cefae9c2e	[log] [tgz]
author	Christopher Faulet <cfaulet@haproxy.com>	Tue Jul 26 19:19:18 2022 +0200
committer	Christopher Faulet <cfaulet@haproxy.com>	Wed Aug 03 09:56:38 2022 +0200
tree	66126def11304d4516db1d832d1bb4cddc158777
parent	ab4b094055fee113e91149426035c9ff7b35b2b0 [diff]

BUG/MEDIUM: peers: limit reconnect attempts of the old process on reload

When peers are configured and HAProxy is reloaded or restarted, a
synchronization is performed between the old process and the new one. To do
so, the old process connects on the new one. If the synchronization fails,
it retries. However, there is no delay and reconnect attempts are not
bounded. Thus, it may loop for a while, consuming all the CPU. Of course, it
is unexpected, but it is possible. For instance, if the local peer is
misconfigured, an infinite loop can be observed if the connection succeeds
but not the synchronization. This prevents the old process to exit, except
if "hard-stop-after" option is set.

To fix the bug, the reconnect is delayed. The local peer already has a
expiration date to delay the reconnects. But it was not used on stopping
mode. So we use it not. Thanks to the previous fix, the reconnect timeout is
shorter in this case (500ms against 5s on running mode). In addition, we
also use the peers resync expiration date to not infinitely retries. It is
accurate because the new process, on its side, use this timeout to switch
from a local resync to a remote resync.

This patch depends on "MINOR: peers: Use a dedicated reconnect timeout when
stopping the local peer". It fixes the issue #1799. It should be backported
as far as 2.0.

src/peers.c[diff]

1 file changed

tree: 66126def11304d4516db1d832d1bb4cddc158777