facebookresearch/fairscale
Question in offload.py: Moving activation to CPU does NOT reduce GPU memory.
Ouverte
#948 ouverte le 4 mars 2022
bughelp wantedoffload_modeltriaged
Métriques du dépôt
- Stars
- (3 411 étoiles)
- Métriques de merge PR
- (Métriques PR en attente)
Description
I use my cuda_active_bytes function to measure the GPU memory before and after the code line below. I find moving activation to CPU does NOT reduce GPU memory.
https://github.com/facebookresearch/fairscale/blob/9f347f373e32ee5cad11a40b70b8e28a74b5e2d4/fairscale/experimental/nn/offload.py#L524
def cuda_active_bytes():
torch.cuda.synchronize()
stats = torch.cuda.memory_stats()
current_active_byte = stats['active_bytes.all.current']
return current_active_byte
So actually all the activations generated by forward is still in GPU memory? If so, I think the code line above is redundant.