Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Clarity on training data for each of the codegen versions

オープン
#76 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
20/100
issue の種類
ドキュメント
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
machine-learning

調査の方向性

Start by checking the "lessons learnt from codegen2" paper and the available documentation for CodeGen2 and CodeGen2.5 base models. Verify whether their training data included ThePile and TheStarCoder, and whether a model at or below 7B used both datasets. Done means documenting a source-backed answer.

索引モデルが issue の本文から書いたものです。

説明

In the "lessons learnt from codegen2" paper, it's discussed that data mix of pile and thestarcoder data is a better choice to undertake if enough compute is available, but it's not clear if codegen2 or codegen2.5 (base models not instruct models) were trained with natural language data like ThePile etc. Is there any small model <=7B which is trained on both ThePile and TheStarCoder data?

主要言語
Python
スター
5.2k
フォーク
420
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

salesforce/CodeGen のほかの issue

salesforce/CodeGen の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。