A question about the detail of data preprocessing
还没有人认领这个 Issue。
评估
- 难度
- 1/5
- 预计耗时
- 1 小时以内
- 新手友好度
- 25/100
- Issue 类型
- 文档
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- python
调研方向
Start with preprocess/1_split_raw.py at line 33 and trace how train.txt is read or produced. Document what data belongs in train.txt and whether it contains all training code; done when the preprocessing input and split expectations are clear.
由索引模型根据 Issue 内容生成。
描述
Hello!
I would like to finetune the model, and during the part of data preprocessing. I saw that in line 33 of the file https://github.com/salesforce/jaxformer/blob/main/preprocess/1_split_raw.py, the code is args.data_bucket_path = '/tmp/dataset_v1/ 0_raw/train.txt'.
I would like to know what kind of data is in the file train.txt? Is all the code data to be trained put into this train.txt file?
- 主要语言
- Python
- 星标
- 5.2k
- 派生
- 420
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
salesforce/CodeGen 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 45/100
salesforce/CodeGen#107 ·
-
难度 5/5 一周以上 新手友好度 20/100
salesforce/CodeGen#104 ·
-
难度 5/5 一周以上 新手友好度 25/100
salesforce/CodeGen#101 ·
-
难度 1/5 1 小时以内 新手友好度 48/100
salesforce/CodeGen#95 ·
-
难度 3/5 1-2 天 新手友好度 35/100
salesforce/CodeGen#94 · 8 条评论 ·
查看 salesforce/CodeGen 的全部 Issue
相似的 Issue
-
bug server
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
sportsdataverse/sportsdataverse-py#641 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
googleapis/google-cloud-python#18532 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复