share_keys() marks outbound Megolm session as shared despite failed Olm session establishment — silent key loss, no recovery

Open
#186 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
30/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
python
Domain
cryptography

Research direction

Start by tracing share_keys() through the crypto manager and the crypto_megolm_outbound_session state described in the issue, using the "No one-time keys nor device keys got when trying to share keys" warning as the failure point. Determine how a failed Olm establishment leads to shared=1 and define whether completion requires retrying pending devices, retaining an unshared session, or another recovery path.

Written by the indexing model from the issue text.

Description

Summary

In mautrix-python 0.21.1, when the crypto manager shares a newly-created outbound Megolm session via share_keys(), a failure to establish an Olm session with a recipient device (e.g. the device has no claimable one-time keys) does not prevent the outbound session from being marked as shared. The share failure is only visible as a mau.crypto logger warning ("No one-time keys nor device keys got when trying to share keys"), which is easy to miss — from that point on, every message encrypted under that outbound session is undecryptable for the affected device, with no room-key request sent by the client and no built-in recovery path.

Observed behaviour (production, 4 episodes Jul–Sep 2026)

Gateway-style bot (the Hermes agent's Matrix platform adapter, which uses mautrix-python for E2EE) running against matrix.org, in encrypted rooms with a user holding 3 active devices:

  1. Process restart → new outbound Megolm session created for the room.
  2. During the initial share_keys(), one or more devices have no available one-time keys. Pairs of warnings are logged:
    WARNING mau.crypto: No one-time keys nor device keys got when trying to share keys
    
    (two warnings per retry — observed 2–6 warnings per episode)
  3. The outbound session is nevertheless marked as shared (shared=1 in crypto_megolm_outbound_session) and all subsequent room messages are encrypted under it.
  4. Affected devices show m.unable_to_decrypt for every later event in the room. The sessions created before the restart remain readable — only the new outbound session's traffic is lost for the omitted device.
  5. The state persists until the outbound Megolm session is discarded and a new one is created after the affected devices have uploaded fresh one-time keys.

Traces from a live episode (2026-09-01, UTC):

01:54:10 WARNING mau.crypto: No one-time keys nor device keys got when trying to share keys
01:54:11 WARNING mau.crypto: No one-time keys nor device keys got when trying to share keys
01:57:37 WARNING mau.crypto: No one-time keys nor device keys got when trying to share keys
...
15:36:32  (new outbound Megolm session created, shared=1, message_count>0)
16:26:37 WARNING mau.client.crypto: Failed to decrypt $...: Failed to decrypt megolm event: no session with given ID <id> found
Expected behaviour

Either (a) the outbound session should not be marked shared while any recipient device failed to receive the key, with the share retried when the device comes back online / uploads OTKs, or (b) the omitted device's client should issue an m.room_key_request that the sender fulfils, per the standard key-recovery path. Today neither happens: the failure is silent at the protocol level, so recovery depends entirely on application-level workarounds.

Suggested directions
  • Return a per-device success/failure result from the share step and let the caller decide whether to keep or rotate the session.
  • Track "pending share" devices on the outbound session and retry the Olm-transport of the room key when the device reappears (OTK claim succeeds), instead of only warning.
  • On decrypt failure with OLM_PREKEY missing, consider prompting a room-key request from the affected side (this may be more natural at the client layer, but a library hook would help).
Workaround used (application level)

Discard the outbound Megolm session row while the process is restarting (the session is cached in memory by the running crypto manager — deleting the DB row while running gets it resurrected from cache), ensure the user's devices have uploaded fresh one-time keys (opening the client suffices), then have any member send a message so a fresh session is created and shared. Verified working, but it requires application-level surgery on the crypto store and a coordinated restart — not something every deployment can do.

Environment
  • mautrix-python 0.21.1 (latest release as of 2026-09-01)
  • Python 3.13, Linux, matrix.org homeserver
  • Bot account with cross-signing verified, store version 10
  • Affected clients: Element X (iOS/iPadOS), Element macOS — i.e. current, maintained clients
Dominant language
Python
Stars
249
Forks
84
PR merge metrics
No merged PRs in 30d

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from mautrix/python

All issues in mautrix/python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.