Discussion: parallelism for borgstore
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Da chiarire
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- backend, performance
Direzione di ricerca
Inizia con il controllo dell’annidamento intorno a src/borgstore/store.py alla riga 208 e confrontalo con l’implementazione parallela proposta nel backend s3.py collegato. Traccia il modo in cui le letture e le scritture di store vengono coordinate, quindi determina quale delle tre proposte —annidamento opzionale, prefetching del backend o un thread pool condiviso— può essere specificata e testata senza compromettere la coerenza.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hello!
Since borgstore allows borg to work with remote repositories, I believe borg/borgstore needs to be able to handle parallel actions to improve the performance when working with remote repositories due to the inherent latency that comes with anything that is not local.
From my tests, I see that with one exception, the backup process can be helped a lot by processing all writes in parallel at store level.
For storing information the main bottleneck is here:
https://github.com/borgbackup/borgstore/blob/a38a1a7c655eaa25dd4e650b9a4deeca4b275446/src/borgstore/store.py#L208
Due to nesting level changes this efectively locks you down to 1 get / 1 put per chunk, sequencially.
By getting rid of the nesting check, I was able to parallelize all the sequencial writes effectively. Doing this, the storage repository is no longer the bottleneck for making backups.
The code I used is this:
https://github.com/alexandru-bagu/borgstore/blob/parallel-s3-store/src/borgstore/backends/s3.py
The idea behind the implementation is that all writes can be parallelized but if you want to read, nothing else is allowed to happen at the same time. This can obviously be improved by using a read-write queue, any reads at any time, any writes at any time, no read and writes at the same time to make sure that the store is consistent.
At the end of this, I am able to easily upload 400 Mbps to S3. The issue now is the extract, because the operations are all sequencial. This is even more of an issue since borg itself is not multi-threaded. My extract network bandwidth is at most 40 Mbps because it is reading one chunk at a time.
One solution for this would be to allow the backend to handle parallelism on its own by providing an ordered list of files that will be requested at some point. That way the backend can prefetch all chunks it can (based on its own logic) and just return them when load is called.
TLDR:
- Can we make nesting a configuration option so that we can effectively skip the nesting checks if we want to? Or make it optional in general?
- Can we update the backend implementation and provide something like an iterator with all the chunks that will be downloaded? Something like a method that each store can implement (it should be optional): prefetch(paths: iterator)
- Can the borgstore have its own thread-pool for parallelism that any backend can make use of? This way there is a structure that is to be followed for parallelism, instead of each backend having its own implementation.
*Cygwin has issues with stopping threads however, if a thread-pool is used this can be non issue because the thread-pool can be disposed only at the end of the execution. This generates a set of warnings/errors but they can be ignored.
Example of warnings (it's like ~1000 lines of this warning at the end):
0 [] python3.9 721 sig_send: error sending signal 11, pid 721, pipe handle 0x15C, nb 0, packsize 192, Win32 error 6
2912 [] python3.9 721 sig_send: error sending signal 11, pid 721, pipe handle 0x15C, nb 0, packsize 192, Win32 error 6
3120 [] python3.9 721 sig_send: error sending signal 11, pid 721, pipe handle 0x15C, nb 0, packsize 192, Win32 error 6
- Lingua principale
- Python
- Stelle
- 26
- Fork
- 9
- Merge medio
- 1h 45m
- PR unite (30g)
- 17
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di borgbackup/borgstore
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
borgbackup/borgstore#234 ·
I maintainer di solito rispondono entro 1 giorno
-
Let backends declare thread_safe and allow concurrent backend calls in Store (follow-up to #206)Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
borgbackup/borgstore#208 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
borgbackup/borgstore#201 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
stats: cache_disabled is False even when no cache backend is configuredForse già presa @MuhammadBilal64 l’ha presa 4 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
borgbackup/borgstore#199 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
write-back cacheApertaenhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
borgbackup/borgstore#169 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di borgbackup/borgstore
Issue simili
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
awslabs/visual-asset-management-system#414 ·
I maintainer di solito rispondono entro 1 giorno
-
bug v1 v2
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
modelcontextprotocol/python-sdk#3670 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
aicell-lab/bioengine#232 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
modelscope/evalscope#1836 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100