Create dataset iarpa_babel_swahili_language_pack
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 62/100
- Issue 类型
- 功能
- 描述清晰度
- 描述清楚
- 活跃度
- 停滞
调研方向
从此 issue 中请求的数据集元数据以及所引用的 IARPA Babel Swahili Language Pack 源开始。按照仓库现有的数据集条目约定,将条目添加为 iarpa_babel_swahili_language_pack.json;当数据集包含其来源、可用性、许可、语言和媒体详细信息时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
- uid: iarpa_babel_swahili_language_pack
- type: processed
- description:
- name: IARPA Babel Swahili Language Pack
- description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
- homepage: https://doi.org/10.35111/afrp-a637
- validated: True
- languages:
- language_names:
- Niger-Congo
- Swahili
- language_comments:
- language_locations:
- Eastern Africa
- Kenya
- validated: False
- language_names:
- custodian:
- name:
- in_catalogue: linguistic_data_consortium_ldc
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - after signing a user agreement
- download_url: https://doi.org/10.35111/afrp-a637
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Multiple licenses:
- license_properties:
- multiple licenses
- research use
- non-commercial use
- copyright - all rights reserved
- license_list:
- other: Other license
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- procurement:
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- audiovisual
- text_format:
- audiovisual_format:
- .WAV
- image_format:
- database_format:
- text_is_transcribed:
- instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
- instance_count: 10K<n<100K
- instance_size: 100<n<10,000
- validated: False
- category:
- fname: iarpa_babel_swahili_language_pack.json
- 主要语言
- HTML
- 星标
- 91
- 派生
- 47
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
bigscience-workshop/data_tooling 的其他 Issue
-
data catalog
难度 1/5 1 小时以内 新手友好度 82/100
-
data catalog
难度 1/5 1 小时以内 新手友好度 62/100
-
data catalog
难度 2/5 1-3 小时 新手友好度 68/100
-
data catalog
难度 1/5 1 小时以内 新手友好度 72/100
-
data catalog
难度 2/5 1-3 小时 新手友好度 82/100
查看 bigscience-workshop/data_tooling 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 86/100
-
难度 2/5 1-3 小时 新手友好度 68/100
cisagov/cyhy-reports#149 · 3 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 88/100
-
难度 2/5 1-3 小时 新手友好度 78/100
-
难度 1/5 1 小时以内 新手友好度 90/100
open-compass/VLMEvalKit#1698 ·