kubeflow/sdk

[SDK] Snapshot users' workspace into distributed TrainJob workload

Open

#48 建立於 2024年12月10日

在 GitHub 查看
 (22 留言) (6 反應) (0 負責人)Python (196 fork)auto 404
help wantedkind/featurelifecycle/frozen

倉庫指標

Star
 (124 star)
PR 合併指標
 (PR 指標待抓取)

描述

What you would like to be added?

As we discussed earlier, we want to design an approach to snapshot users' workspace into TrainJob (e.g. distributed ML workload): https://github.com/kubeflow/training-operator/pull/2324#discussion_r1862719941. To achieve this, we plan to generate a unique TrainJob ID before submitting it to the Kubernetes control plane.

During the KubeCon 2024 demo, we demonstrated how workspace snapshotting might work: https://youtu.be/Lgy4ir1AhYw?t=458. In this demo, we pushed Python code files into S3 and then loaded them into TrainJob using initContainers.

However, we can consider various approaches, for instance:

  • Using distributed cache.
  • Using kubectl cp.

Why is this needed?

This should streamline Data Scientists user experience while working with Kubeflow Training Python SDK.

Love this feature?

Give it a 👍 We prioritize the features with most 👍

貢獻者指南