Create dataset aaj_tak

Open Beginner friendly
#98 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
68/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Stale
Tech stack
json
Domain
data

Research direction

Start by comparing the requested aaj_tak.json entry with existing dataset metadata entries in the repository. Add the Aaj Tak fields shown in the issue, preserve the schema and values, and verify that the new JSON is valid and included wherever dataset entries are catalogued.

Written by the indexing model from the issue text.

Description

data catalog need custodian permission
  • uid: aaj_tak
  • type: primary
  • description:
  • languages:
    • language_names:
      • Indic
      • Hindi
    • language_comments:
    • language_locations:
      • Southern Asia
      • India
    • validated: False
  • custodian:
    • name: Aaj Tak
    • in_catalogue:
    • type: A commercial entity
    • location: India
    • contact_name: Aaj Tak
    • contact_email: info@aajtak.com
    • contact_submitter: False
    • additional:
    • validated: False
  • availability:
    • procurement:
      • for_download: No - we would need to spontaneously reach out to the current owners/custodians
      • download_url:
      • download_email: info@aajtak.com
    • licensing:
      • has_licenses: Yes
      • license_text:
      • license_properties:
        • copyright - all rights reserved
      • license_list:
    • pii:
      • has_pii: Yes
      • generic_pii_likely:
      • generic_pii_list:
      • numeric_pii_likely:
      • numeric_pii_list:
      • sensitive_pii_likely:
      • sensitive_pii_list:
      • no_pii_justification_class:
      • no_pii_justification_text:
    • validated: False
  • source_category:
    • category_type: website
    • category_web: news or magazine website
    • category_media:
    • validated: False
  • media:
    • category:
      • text
    • text_format:
      • .TXT
      • .PAGES
    • audiovisual_format:
    • image_format:
      • .JPG
    • database_format:
    • text_is_transcribed: Yes - image
    • instance_type:
    • instance_count:
    • instance_size:
    • validated: False
  • fname: aaj_tak.json
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from bigscience-workshop/data_tooling

All issues in bigscience-workshop/data_tooling

Similar issues

More Data Engineering issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.