Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Issue with storing a parsetree

未关闭
#150 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
20/100
Issue 类型
功能
描述清晰度
需要澄清
活跃度
停滞
技术栈
python
领域
data

调研方向

Issue 没有指出任何 repository 文件或测试。请从 parsetree(...) 调用以及 Text 和 Sentence 类开始,然后检查它们的内容是如何表示的,以及是否有现有的导出接口文档。要视为完成,需要由 maintainer 定义输出格式并进行相应的验证,而 issue 没有对此作出规定。

由索引模型根据 Issue 内容生成。

描述

Hello !

Painful issue right here.

I have built a parse tree quite simply on a large volume of texts like :

s = parsetree(string, relations=True, lemmata=True)

with s being of the type : <class 'pattern.text.tree.Text'>

If I do a pprint(s) I get a very clean data structure like :


             WORD   TAG    CHUNK   ROLE   ID     PNP    LEMMA               

@MAP_Information   NN     NP      -      -      -      @map_information   
               et   CC     -       -      -      -      et                  
          pendant   IN     PP      -      -      PNP    pendant             
               ce   PRP    NP      SBJ    1      PNP    ce                  
            temps   NN     NP ^    SBJ    1      PNP    temps               

Which is want I want ! So I would like to store the exact same data structure to any file like a DataFrame, a CSV, plain text... for better readability and user-friendliness.

However this is not possible since it all the outputs belong to <class 'pattern.text.tree.Text'> or <class 'pattern.text.tree.Sentence'> etc... classes, which make them painful to use.

For example I cannot :

encode my object to utf-8 for exporting :

s = s.encode('utf8')

AttributeError: 'Text' object has no attribute 'encode'

Export my object as a text file :

with open("D:\\Testpprint4.txt", 'w') as export :
    for sentence in s :
        export.write(sentence)

TypeError: expected a string or other character buffer object

Put my object to a dataframe :

`pd = pandas.DataFrame(sentence)

PandasError : DataFrame constructor not properly called!`

Or use pprint.pformat to store it as a csv since it does not deal with utf-8.

Would you have a solution ?

Thanks,

主要语言
Python
星标
8.9k
派生
1.6k
PR 合并指标
30 天内没有已合并 PR

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

clips/pattern 的其他 Issue

查看 clips/pattern 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。