rapidsai/cudf

[BUG] `cudf.read_text` throws an exception when reading a host buffer

已关闭

#13,734 创建于 2023年7月22日

 (2 条评论) (0 个反应) (0 位负责人)C++ (735 个派生)batch import
0 - BacklogPythonbugcuIOgood first issue

仓库指标

星标
 (6,000 个星标)
PR 合并指标
 (平均合并 17天 21小时) (30 天内合并 230 个 PR)

描述

Describe the bug cudf.read_text throws an exception when reading a host buffer, unlike the CSV, Parquet, ORC and JSON readers. Hopefully it is straightforward to enable cudf.read_text to read from host buffers.

Steps/Code to reproduce bug Here is a small code snippet that shows the exception:

from io import BytesIO
import cudf

df = cudf.DataFrame({'a':['aaaa','bbbb']})

buf = BytesIO()
df.to_csv(buf, index=False)
df2 = cudf.read_csv(buf) # this is ok
print(df2)

buf = BytesIO()
df.to_csv(buf, index=False)
df2 = cudf.read_text(buf, delimiter='\n') # this crashes
print(df2)

Expected behavior I expect that read_text would support host buffers correctly in the python layer. "multibyte_split" works fine with HOST_BUFFER data source in the libcudf benchmarks.

Environment overview (please complete the following information) nightly docker image from 23.08, A100 DGX workstation

贡献者指南