Create dataset iarpa_babel_swahili_language_pack

Open Beginner friendly
#128 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
62/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Stale

Research direction

Start with the requested dataset metadata in this issue and the referenced IARPA Babel Swahili Language Pack source. Add the entry as iarpa_babel_swahili_language_pack.json using the repository’s existing dataset-entry conventions; done means the dataset is represented with its source, availability, licensing, language, and media details.

Written by the indexing model from the issue text.

Description

data catalog need custodian permission need data sourcing feedback
  • uid: iarpa_babel_swahili_language_pack
  • type: processed
  • description:
    • name: IARPA Babel Swahili Language Pack
    • description: Swahili ASR Dataset. Official description says "IARPA Babel Swahili Language Pack IARPA-babel202b-v1.0d was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 350 hours of Swahili conversational and scripted telephone speech collected from 2012-2014 along with corresponding transcripts."
    • homepage: https://doi.org/10.35111/afrp-a637
    • validated: True
  • languages:
    • language_names:
      • Niger-Congo
      • Swahili
    • language_comments:
    • language_locations:
      • Eastern Africa
      • Kenya
    • validated: False
  • custodian:
    • name:
    • in_catalogue: linguistic_data_consortium_ldc
    • type:
    • location:
    • contact_name:
    • contact_email:
    • contact_submitter: False
    • additional:
    • validated: False
  • availability:
  • processed_from_primary:
    • from_primary: Taken from primary source
    • primary_availability: No - the dataset curators kept the source data secret
    • primary_license:
    • primary_types:
    • validated: False
  • media:
    • category:
      • audiovisual
    • text_format:
    • audiovisual_format:
      • .WAV
    • image_format:
    • database_format:
    • text_is_transcribed:
    • instance_type: 350 hours of telephone conversations, no clue how long each file is. Estimating 15 minutes each, and 350 words per minute?
    • instance_count: 10K<n<100K
    • instance_size: 100<n<10,000
    • validated: False
  • fname: iarpa_babel_swahili_language_pack.json
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from bigscience-workshop/data_tooling

All issues in bigscience-workshop/data_tooling

Similar issues

More Data Engineering issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.