LoRA weight merging giving torch distributed error on single-node single-gpu
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 25/100
- Issue 類型
- 缺陷
- 描述清晰度
- 需要釐清
- 活躍度
- 停滯
- 技術堆疊
- jupyter-notebook, python, pytorch
研究方向
從 finetuning/Gemma/lora.ipynb 和 /opt/NeMo/scripts/nlp_language_modeling/merge_lora_weights/merge.py 開始,然後使用列出的 overrides 重現單節點、單 GPU 的 merge。調查為什麼在還原模型時,分散式初始化會嘗試繫結連接埠 53747。對於所回報的設定,如果 LoRA 和模型權重能夠成功 merge 且不出現 socket error,即表示完成。
由索引模型根據 Issue 內容生成。
描述
I am running this notebook. However, when I try to merge LoRA and model weights before exporting to TensorRTLLM (python /opt/NeMo/scripts/nlp_language_modeling/merge_lora_weights/merge.py). I received the following error:
Initializing distributed: GLOBAL_RANK: 0, MEMBER: 1/1
[W socket.cpp:464] [c10d] The server socket has failed to bind to [::]:53747 (errno: 98 - Address already in use).
[W socket.cpp:464] [c10d] The server socket has failed to bind to ?UNKNOWN? (errno: 98 - Address already in use).
[E socket.cpp:500] [c10d] The server socket has failed to listen on any local network address.
Error executing job with overrides: ['trainer.accelerator=gpu', 'tensor_model_parallel_size=1', 'pipeline_model_parallel_size=1', 'gpt_model_file=gemma_2b_pt.nemo', 'lora_model_path=nemo_experiments/gemma_lora_pubmedqa/checkpoints/gemma_lora_pubmedqa.nemo', 'merged_model_path=gemma_lora_pubmedqa_merged.nemo']
Traceback (most recent call last):
File "/opt/NeMo/scripts/nlp_language_modeling/merge_lora_weights/merge.py", line 171, in main
model = MegatronGPTModel.restore_from(
File "/usr/local/lib/python3.10/dist-packages/nemo/collections/nlp/models/nlp_model.py", line 478, in restore_from
return super().restore_from(
File "/usr/local/lib/python3.10/dist-packages/nemo/core/classes/modelPT.py", line 468, in restore_from
instance = cls._save_restore_connector.restore_from(
File "/usr/local/lib/python3.10/dist-packages/nemo/collections/nlp/parts/nlp_overrides.py", line 1306, in restore_from
trainer.strategy.setup_environment()
File "/usr/local/lib/python3.10/dist-packages/pytorch_lightning/strategies/ddp.py", line 154, in setup_environment
self.setup_distributed()
File "/usr/local/lib/python3.10/dist-packages/nemo/collections/nlp/parts/nlp_overrides.py", line 244, in setup_distributed
super().setup_distributed()
File "/usr/local/lib/python3.10/dist-packages/pytorch_lightning/strategies/ddp.py", line 203, in setup_distributed
_init_dist_connection(self.cluster_environment, self._process_group_backend, timeout=self._timeout)
File "/usr/local/lib/python3.10/dist-packages/lightning_fabric/utilities/distributed.py", line 297, in _init_dist_connection
torch.distributed.init_process_group(torch_distributed_backend, rank=global_rank, world_size=world_size, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/c10d_logger.py", line 86, in wrapper
func_return = func(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/distributed_c10d.py", line 1172, in init_process_group
store, rank, world_size = next(rendezvous_iterator)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/rendezvous.py", line 244, in _env_rendezvous_handler
store = _create_c10d_store(master_addr, master_port, rank, world_size, timeout, use_libuv)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/rendezvous.py", line 172, in _create_c10d_store
return TCPStore(
torch.distributed.DistNetworkError: The server socket has failed to listen on any local network address. The server socket has failed to bind to [::]:53747 (errno: 98 - Address already in use). The server socket has failed to bind to ?UNKNOWN? (errno: 98 - Address already in use).
Setup Information:
torch: 2.2.0a0+81ea7a4
nemo: 2.0
Container: nvcr.io/nvidia/nemo:24.01.gemma
- 主要語言
- Jupyter Notebook
- 星號
- 4.2k
- 分支
- 1.1k
- 平均合併
- 10 小時 15 分鐘
- 30 天內合併 PR
- 1
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
NVIDIA/GenerativeAIExamples 的其他 Issue
-
Incorrect audit job payload for nemo-microservices/auditor:25.09可能已有人在做 @saichandrapandraju 於 363 天前認領。 未關閉
難度 1/5 1 小時以內 新手友好度 70/100
NVIDIA/GenerativeAIExamples#361 · 2 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 72/100
NVIDIA/GenerativeAIExamples#299 ·
-
難度 3/5 1-2 天 新手友好度 35/100
NVIDIA/GenerativeAIExamples#437 ·
-
hyperlink not working可能已有人在做 @maowiz 於 72 天前認領。 未關閉
難度 1/5 1 小時以內 新手友好度 45/100
NVIDIA/GenerativeAIExamples#400 ·
-
難度 5/5 一週以上 新手友好度 10/100
NVIDIA/GenerativeAIExamples#399 ·
查看 NVIDIA/GenerativeAIExamples 的全部 Issue
相似的 Issue
-
auth manager-contract: serve grant takes an unused serve actor and computes a discarded actors value未關閉area:auth bug
難度 2/5 1-3 小時 新手友好度 88/100
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 72/100
維護者通常 1 天內回覆
-
defect
難度 2/5 1-3 小時 新手友好度 88/100
-
Distributed group_mean casts group ids to float32, losing precision vs non-distributed path可能已有人在做 @OnePunchMonk 於 1 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 82/100
維護者通常 1 天內回覆
-
[Refactor♻️] Rename RoleChangeOutcome to BrokerRoleChangeReport可能已有人在做 @niukanen1 於 1 天前認領。 未關閉Difficulty level/Easy good first issue help wanted refactor♻️ rocketmq-broker crate rust
難度 2/5 1-3 小時 新手友好度 88/100
mxsm/rocketmq-rust#11123 · 1 則留言 · 已指派 1 人 ·
維護者通常 1 天內回覆