Race condition during sync of large projects can block dbsync
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 38/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Stale
- Tech stack
- postgresql, python
- Domain
- databases, distributed-systems
Research direction
Read the version check in dbsync.py's push function and trace the existing pull function's rebase behavior. Test the large-project version-mismatch scenario described in the issue; done means a mismatch can pull the latest changes and retry safely, while a second mismatch still reports an error without requiring --force-init.
Written by the indexing model from the issue text.
Description
Description
A race condition exists in dbsync that can block the synchronization process, especially with large projects that take a long time to download. When dbsync initiates a pull operation, and another client pushes a new version to the Mergin Maps server before the pull is complete, dbsync ends up with an outdated local version of the project.
This leads to a failure in the subsequent push operation, because of a strict version check that ensures the local version matches the server version. The push function raises an error: "There are pending changes on server - need to pull them first.". This creates a loop where dbsync is stuck trying to pull, but each pull is slow and susceptible to the same race condition, requiring manual intervention like --force-init, which can lead to data loss.
Why --force-init is not a solution
Using --force-init is a heavy-handed approach that wipes the local state and re-initializes the synchronization from scratch. This is not a viable solution in a production environment for several reasons:
- Data Loss: If there are changes in the PostgreSQL database that have not been pushed to the Mergin Maps server, a
--force-initwill wipe thebaseandmodifiedschemas and re-create them from the GeoPackage file. This will cause any changes made in the database to be lost. - Manual Intervention: The need for manual intervention defeats the purpose of an automated synchronization daemon.
- Downtime: The re-initialization process can be time-consuming for large projects, leading to extended downtime for the synchronization service.
The problematic version check is located in the push function in dbsync.py:
# dbsync.py in push()
# ...
# check there are no pending changes on server
if server_version != local_version:
raise DbSyncError("There are pending changes on server - need to pull them first.")
Real-world Scenario
- T0:
dbsyncstarts apulloperation for a large project with many photos. The server is at versionv100. The download is expected to take over a minute. - T0 + 30s: A surveyor in the field finishes their work and syncs their mobile client. This creates version
v101on the Mergin Maps server. - T0 + 90s:
dbsynccompletes its download ofv100and applies the changes to the PostgreSQL database. The local project version fordbsyncis nowv100. - T0 + 95s: The
dbsyncdaemon proceeds to thepushstep to sync changes from the database back to Mergin Maps. - Failure: The
pushoperation detects that the server is atv101while the local version isv100. It aborts the push, anddbsyncis effectively blocked.
Proposed Solution
To resolve this, the push function should be made more resilient. Instead of immediately failing upon a version mismatch, it should attempt to resolve the situation automatically by pulling the latest changes.
The proposed solution is to modify the push function in dbsync.py. When a version mismatch is detected, dbsync should:
- Automatically trigger the
pullfunction. The existingpullfunction is capable of handling a rebase of local database changes on top of the incoming server changes. - After the
pullis complete, re-check the version. - If the versions now match, proceed with the
pushoperation. - If the versions still do not match after the automatic pull, then raise an error, as this would indicate a more serious problem that requires manual intervention.
This "pull-and-retry" mechanism would make the synchronization process more robust for projects with long download times and active collaboration, avoiding the need for manual resets.
- Dominant language
- Python
- Stars
- 53
- Forks
- 24
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from MerginMaps/db-sync
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
MerginMaps/db-sync#184 ·
-
Difficulty 3/5 1-2 days Newbie friendliness 38/100
MerginMaps/db-sync#185 · 1 comment ·
-
db-sync enters an infinite drop/recreate loop when geodiff init fails on invalid source geometries Open
Difficulty 4/5 3-5 days Newbie friendliness 42/100
MerginMaps/db-sync#182 ·
-
Difficulty 3/5 1-2 days Newbie friendliness 52/100
MerginMaps/db-sync#181 · 2 comments ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
MerginMaps/db-sync#179 · 3 comments ·
All issues in MerginMaps/db-sync
Similar issues
-
documentation help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
simonw/sqlite-utils#872 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100