Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

A question about the detail of data preprocessing

Đang mở
#69 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
1/5
Thời gian dự kiến
Dưới một giờ
Mức phù hợp với người mới
25/100
Loại issue
Tài liệu
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
python

Hướng nghiên cứu

Start with preprocess/1_split_raw.py at line 33 and trace how train.txt is read or produced. Document what data belongs in train.txt and whether it contains all training code; done when the preprocessing input and split expectations are clear.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Hello!

I would like to finetune the model, and during the part of data preprocessing. I saw that in line 33 of the file https://github.com/salesforce/jaxformer/blob/main/preprocess/1_split_raw.py, the code is args.data_bucket_path = '/tmp/dataset_v1/ 0_raw/train.txt'.

I would like to know what kind of data is in the file train.txt? Is all the code data to be trained put into this train.txt file?

Ngôn ngữ chính
Python
Star
5.2k
Fork
420
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của salesforce/CodeGen

Tất cả issue của salesforce/CodeGen

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.