[REQUEST] More fine-grained distributed strategies for RLHF training
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Estancado
- Stack tecnológico
- python
Línea de trabajo
No se nombran archivos del repositorio, pruebas ni puntos de entrada. Empieza por localizar la canalización de entrenamiento RLHF existente y comparar la ubicación de sus modelos con la del artículo APP enlazado en el issue; se considera terminado cuando se hayan integrado y validado las estrategias propuestas Separation e Interleaving, incluida su ubicación distribuida y su comportamiento durante la etapa de generación.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Is your feature request related to a problem? Please describe.
We find that the generation stage of RLHF pipeline is time-consuming during the current training process. This is because the four models (Actor, Critic, Reward, and Ref) are all colocated on the same devices, utilizing a "Flattening" strategy. This results in that both training and inference runtime are mixed in the current procedure. It disables the training or inference specialized optimization methods. Also, a significant amount of memory is occupied by models, but they are idle in generation stage of actor model. Therefore, instead of collocating these four models on all devices, more fine-grained placement strategy could be utilized.
Describe the solution you'd like
Our team is planning to open-source our implementation of APP (https://arxiv.org/pdf/2312.11819.pdf) and contribute it to the codebase. Specifically, we are proposing two fine-grained model placement strategies:
A Separation strategy that separates the training and inference runtime of the RLHF pipeline with additional shadow models. This enables the adoption of inference-optimized techniques such as vLLM and intra-node tensor parallelism to accelerate the time-cost generation stage. This enables different distributed stragies during the generation stage compared with training stage.
An Interleaving strategy that helps reduce memory redundancy and communication costs in RLHF training by placing models without dependencies on exclusive devices with careful orchestration. For example, inference models like the reward model and reference model could be placed on separate devices. This approach enables the reduction of memory redundancy using the DDP or ZeRO 1-2 by decreasing the scale of participating nodes.
Describe alternatives you've considered
N/A
Additional context
Thank you for sharing the deepspeed-chat with the community! It has been an essential infrastructure, providing an easy-to-use solution for training InstructGPT-like models. Recently, we have made some improvements to further enhance training performance while maintaining the simplicity of usage. These improvements have already been implemented in the RLHF training at Ant Group. In order to share our efforts with the deepspeed-chat community, we would like to integrate our implementation into DeepSpeedExamples codebase.
To facilitate discussions and minimize potential conflicts of interest, we have created this issue to engage in conversations about the proposed modifications. We look forward to collaborating with the community on this matter.
Please feel free to comment here or reach out via email ([email protected]). Thanks!
- Lenguaje dominante
- Python
- Estrellas
- 6.8k
- Forks
- 1.1k
- Merge medio
- 3 d 20 h
- PR fusionados (30 d)
- 5
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de deepspeedai/DeepSpeedExamples
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 52/100
deepspeedai/DeepSpeedExamples#996 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
deepspeedai/DeepSpeedExamples#995 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 52/100
deepspeedai/DeepSpeedExamples#989 ·
-
moe example 404Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
deepspeedai/DeepSpeedExamples#984 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 42/100
deepspeedai/DeepSpeedExamples#979 · 6 comentarios ·
Todos los issues de deepspeedai/DeepSpeedExamples
Issues similares
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Abiertodocumentation priority: low size: S
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Los mantenedores suelen responder en 1 día
-
--csv-bom was never wired up: PR #850 added an unused helper parameter, so #846 is not fixedAbiertobug help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Los mantenedores suelen responder en 1 día
-
Broken link in index.rstAbiertodocumentation
Dificultad 1/5 Menos de una hora Aptitud para principiantes 65/100
ansys/pydpf-core#3547 ·
Los mantenedores suelen responder en 1 día
-
core
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
vectorize-io/hindsight#5457 ·
Los mantenedores suelen responder en 1 día
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recordingPosiblemente ocupada @ktz03 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
volcengine/OpenViking#5806 ·
Los mantenedores suelen responder en 1 día