Create dataset australian_twittersphere
Nobody has claimed this yet.
Assessment
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Newbie friendliness
- 65/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Stale
- Tech stack
- json
- Domain
- data
Research direction
Start by inspecting the dataset entry format and add the supplied Australian Twittersphere record to australian_twittersphere.json. Check that the fields, values, and filename match the surrounding catalog conventions; the work is done when the new dataset entry is accepted by the repository's validation process.
Written by the indexing model from the issue text.
Description
- uid: australian_twittersphere
- type: processed
- description:
- name: Australian Twittersphere
- description: The Australian Twittersphere is a longitudinal, curated collection of tweets from approximately 838,000 Twitter accounts identified as ‘Australian’. The Digital Observatory has maintained reliable, ongoing data collection since early 2018, with approximately 23 million tweets being collected per month. There is also an archive of approximately 2 billion tweets from 2006 to 2016. The Digital Observatory currently collects approximately 37 million tweets per month.
- homepage: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: True
- languages:
- language_names:
- English
- language_comments: Australian English
- language_locations:
- Oceania
- Australia
- validated: False
- language_names:
- custodian:
- name: Digital Observatory of the Queensland University of Technology
- in_catalogue:
- type: A university or research institution
- location: Australia
- contact_name:
- contact_email: digitalobservatory@qut.edu.au
- contact_submitter: False
- additional: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: False
- availability:
- procurement:
- for_download: No - we would need to spontaneously reach out to the current owners/custodians
- download_url:
- download_email: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory/services-and-equipment
- licensing:
- has_licenses: Unclear
- license_text: The data should be able to be used to train models while respecting the rights and wishes of the data creators and custodians, as they were obtained in compliance with Twitter's terms of use.
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: other
- no_pii_justification_text: The data were obtained from Twitter and should have been anonimysed.
- validated: False
- procurement:
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - their documentation/homepage/description is available
- primary_license: Yes - the dataset curators have obtained consent from the source material owners
- primary_types:
- web | social media
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type: post
- instance_count: n>1B
- instance_size: 10<n<100
- validated: False
- category:
- fname: australian_twittersphere.json
- Dominant language
- HTML
- Stars
- 91
- Forks
- 47
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from bigscience-workshop/data_tooling
-
data catalog
Difficulty 1/5 Under an hour Newbie friendliness 82/100
-
data catalog
Difficulty 1/5 Under an hour Newbie friendliness 62/100
-
data catalog
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
data catalog
Difficulty 1/5 Under an hour Newbie friendliness 72/100
-
data catalog
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
All issues in bigscience-workshop/data_tooling
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
-
removeToken
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
cowprotocol/token-lists#1514 · 2 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
anomalyco/models.dev#7670 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
cisagov/cyhy-reports#149 · 3 comments ·
-
remove entry
Difficulty 1/5 Under an hour Newbie friendliness 74/100