Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Chapter 8: Shuffling the DataFrame in newer versions of pandas

オープン
#74 コメント 1 件 リアクション 2 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
55/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
停滞
技術スタック
pandas, python

調査の方向性

235ページのDataFrameシャッフルコードから始め、データセットをCSVにエクスポートする前に、pandas 0.23.4でどのように動作するかを確認してください。結果のCSVが実際にシャッフルされていること、また246-246ページのsentiment-analysisの例が、順序付けられたデータによる誤解を招く精度を報告しなくなっていることを検証してください。

索引モデルが issue の本文から書いたものです。

説明

Just a note in case it's helpful to anyone else - I seemed to be getting 100% accuracy with the on-line sentiment analysis classifier (pages 246-246), but it turned out to be because the code used to shuffle the dataset before exporting it to CSV on page 235 hadn't worked.

In the version of pandas I'm using (0.23.4), it looks like df.index.values is needed in order to get the indexes of a DataFrame as a list. So, this:

df = df.reindex(np.random.permutation(df.index))

now needs to be this:

df = df.reindex(np.random.permutation(df.index.values))

Hope that helps someone!

主要言語
Jupyter Notebook
スター
12.7k
フォーク
4.4k
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

rasbt/python-machine-learning-book のほかの issue

rasbt/python-machine-learning-book の issue をすべて見る

似ている issue

Machine Learning の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。