Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Determining the unit of analysis for the machine learning models

オープン
#26 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
machine-learning

調査の方向性

選択の粒度、コンテキストの保存、カテゴリラベルとバイナリラベルの違い、段落レベルの入力に関する issue の質問から始めます。信頼できるコントリビューターがどのようにコンテンツを選択するのかを理解するため、リンクされた Chrome 拡張機能を確認します。ML モデルのトレーニングに向けた分析単位、保持するコンテキスト、ラベリングのアプローチについて合意できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

design-thinking documentation Machine Learning question stale

For the chrome extension, we need to decide what is the unit or type of selections trusted contributors can select when identifying racially biased content. This will be used to train the ML models and is important to help users understand why content is potentially racially biased and to offer up alternatives.

Considerations

  • If content is considered racially biased, context will be a factor. How can we store more data around the selected word / phrase to provide more context for the ML models?
  • At what granularity should the trusted contributors be expected to flag content? (Categorically vs binary?)
  • Consider the input as a paragraph/group of text and all kind of tokenizations can be used to finally preprocess
主要言語
Jupyter Notebook
スター
8
フォーク
8
PR マージ指標
30日以内にマージされた PR はありません

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Call-for-Code-for-Racial-Justice/TakeTwo-DataScience のほかの issue

Call-for-Code-for-Racial-Justice/TakeTwo-DataScience の issue をすべて見る

似ている issue

Data Engineering の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。