Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Chapter 8: Shuffling the DataFrame in newer versions of pandas

未关闭
#74 1 条评论 2 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
55/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
停滞
技术栈
pandas, python

调研方向

从第 235 页的 DataFrame shuffle 代码开始,在数据集导出为 CSV 之前,检查它在 pandas 0.23.4 下的行为。确认生成的 CSV 确实经过了 shuffle,并确认第 246-246 页的 sentiment-analysis 示例不再因有序数据而报告误导性的准确率。

由索引模型根据 Issue 内容生成。

描述

Just a note in case it's helpful to anyone else - I seemed to be getting 100% accuracy with the on-line sentiment analysis classifier (pages 246-246), but it turned out to be because the code used to shuffle the dataset before exporting it to CSV on page 235 hadn't worked.

In the version of pandas I'm using (0.23.4), it looks like df.index.values is needed in order to get the indexes of a DataFrame as a list. So, this:

df = df.reindex(np.random.permutation(df.index))

now needs to be this:

df = df.reindex(np.random.permutation(df.index.values))

Hope that helps someone!

主要语言
Jupyter Notebook
星标
12.7k
派生
4.4k
PR 合并指标
30 天内没有已合并 PR

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

rasbt/python-machine-learning-book 的其他 Issue

查看 rasbt/python-machine-learning-book 的全部 Issue

相似的 Issue

更多 Machine Learning Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。