Create dataset iarpa_babel_swahili_language_pack

未关闭 适合新手
#128 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
62/100
Issue 类型
功能
描述清晰度
描述清楚
活跃度
停滞

调研方向

从此 issue 中请求的数据集元数据以及所引用的 IARPA Babel Swahili Language Pack 源开始。按照仓库现有的数据集条目约定,将条目添加为 iarpa_babel_swahili_language_pack.json;当数据集包含其来源、可用性、许可、语言和媒体详细信息时,即表示完成。

由索引模型根据 Issue 内容生成。

描述

data catalog need custodian permission need data sourcing feedback
  • uid: iarpa_babel_swahili_language_pack
  • type: processed
  • description:
    • name: IARPA Babel Swahili Language Pack
    • description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
    • homepage: https://doi.org/10.35111/afrp-a637
    • validated: True
  • languages:
    • language_names:
      • Niger-Congo
      • Swahili
    • language_comments:
    • language_locations:
      • Eastern Africa
      • Kenya
    • validated: False
  • custodian:
    • name:
    • in_catalogue: linguistic_data_consortium_ldc
    • type:
    • location:
    • contact_name:
    • contact_email:
    • contact_submitter: False
    • additional:
    • validated: False
  • availability:
  • processed_from_primary:
    • from_primary: Taken from primary source
    • primary_availability: No - the dataset curators kept the source data secret
    • primary_license:
    • primary_types:
    • validated: False
  • media:
    • category:
      • audiovisual
    • text_format:
    • audiovisual_format:
      • .WAV
    • image_format:
    • database_format:
    • text_is_transcribed:
    • instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
    • instance_count: 10K<n<100K
    • instance_size: 100<n<10,000
    • validated: False
  • fname: iarpa_babel_swahili_language_pack.json
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

bigscience-workshop/data_tooling 的其他 Issue

查看 bigscience-workshop/data_tooling 的全部 Issue

相似的 Issue

更多 Data Engineering Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。