Skip to content

Persist peer address confirmation and learning mode in Redis - #2184

Open
baseballyama wants to merge 2 commits into
sipwise:masterfrom
baseballyama:redis-endpoint-learning
Open

baseballyama wants to merge 2 commits into
sipwise:masterfrom
baseballyama:redis-endpoint-learning

Conversation

@baseballyama

@baseballyama baseballyama commented Sep 28, 2026 •

Copy link
Copy Markdown

close: #2182

A call restored from Redis re-learns its peer address, and the learning mode requested in the offer is lost (#2182). This writes the sfd's confirmed flag and the stream's el_flags and learned_endpoint to the record and reads them back on restore. Records without the new keys restore as before.

The second commit is needed once the confirmation survives a takeover. The first packet that arrives after active cleared FOREIGN_MEDIA even if it was then dropped by strict-source, and a call that had been restored longer than timeout ago was closed before its real media arrived.

Tested with two instances using keyspace notifications and active-switchover, kernel forwarding. 100 calls with endpoint-learning-off and strict-source carried genuine RTP, and spoofed RTP was sent to every media port while switching over:

calls taken over by the spoofed source calls with genuine media at the end
master 73 27
first commit only 0 99 (one closed by the timeout)
this PR 0 100

The Redis tests now expect the new keys, and the keyspace subscription test checks that a restored stream comes back confirmed.

This PR and #2185 touch the same test files but merge cleanly with each other.


This PR was prepared with the assistance of Claude Code (Claude Opus 5.5). The measurements are from our own test environment.

A call restored from Redis starts with every sfd unconfirmed and every
stream's el_flags at zero, which is EL_DELAYED, regardless of what the
offer asked for or what --endpoint-learning is set to. After a
keyspace-notification failover, all adopted calls therefore re-learn
their peer at the same time, and strict-source has nothing to compare
against until they have.

Write the sfd's confirmed flag and the stream's el_flags and
learned_endpoint to the record, and read them back on restore. The two
stream fields were so far only part of the rollback snapshot. Move them
into the shared stream encoder and decoder, so that both use the same
definition. Records without the new keys restore as before.

Fixes sipwise#2182
After a takeover, FOREIGN_MEDIA holds off the media timeout until the
call receives media. The flag is cleared by the first packet that
arrives on any of the call's sockets, before the packet is checked, so
a packet that is then dropped, for example because of strict-source,
ends the grace period without refreshing the stream's last packet time.
A call that was restored longer than the timeout ago is then closed on
the next timer run, before its real media has had a chance to arrive.

With the peer address confirmation now restored from Redis, this is
what happens to an adopted strict-source call when a stray packet
arrives ahead of the peer's first one.

Clear the flag only once a packet has passed the address check. The
timer still clears it for kernel-forwarded media, and the grace period
stays bounded by the timeout after the takeover.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Endpoint-learning state (confirmed, el_flags) is not restored from Redis, so every adopted call re-learns its peer after failover

1 participant