Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Clarity on training data for each of the codegen versions

Đang mở
#76 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
20/100
Loại issue
Tài liệu
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
machine-learning

Hướng nghiên cứu

Start by checking the "lessons learnt from codegen2" paper and the available documentation for CodeGen2 and CodeGen2.5 base models. Verify whether their training data included ThePile and TheStarCoder, and whether a model at or below 7B used both datasets. Done means documenting a source-backed answer.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

In the "lessons learnt from codegen2" paper, it's discussed that data mix of pile and thestarcoder data is a better choice to undertake if enough compute is available, but it's not clear if codegen2 or codegen2.5 (base models not instruct models) were trained with natural language data like ThePile etc. Is there any small model <=7B which is trained on both ThePile and TheStarCoder data?

Ngôn ngữ chính
Python
Star
5.2k
Fork
420
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của salesforce/CodeGen

Tất cả issue của salesforce/CodeGen

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.