Create dataset AOC
Abierto
@apergo-ai ya está trabajando en esto.
Desde el 10/1/2022.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
data catalog
need data sourcing feedback
Source: Masader Project
- uid: arabic_online_commentary
- entry: https://arbml.github.io/masader/card.html?39
- Link: https://github.com/sjeblee/AOC
- License : unknown
- Year: 2011
- Language: ar
- Dialect: other
- Domain: news articles
- Form: text
- Collection Style: crawling and annotation(other)
- Description: a 52M-word monolingual dataset rich in dialectal content
- Volume: 108,000
- Unit: sentences
- Ethical Risks: Low
- Provider: Johns Hopkins University
- Derived From:
- Paper Title: The Arabic Online Commentary Dataset: an Annotated Dataset of Informal Arabic with High Dialectal Content
- Paper Link: https://aclanthology.org/P11-2007.pdf
- Script: Arab
- Tokenized: No
- Host: GitHub
- Access: Free
- Cost:
- Test Split: No
- Tasks: dialect identification
- Evaluation Set?:
- Venue Title: ACL
- Citations: 147
- Venue Type: conference
- Venue Name: Assofications of computation linguisitcs
- authors: Omar Zaidan,Chris Callison-Burch
- affiliations: ,
- abstract: The written form of Arabic, Modern Standard Arabic (MSA), differs quite a bit from the spoken dialects of Arabic, which are the true "native" languages of Arabic speakers used in daily life. However, due to MSA's prevalence in written form, almost all Arabic datasets have predominantly MSA content. We present the Arabic Online Commentary Dataset, a 52M-word monolingual dataset rich in dialectal content, and we describe our long-term annotation effort to identify the dialect level (and dialect itself) in each sentence of the dataset. So far, we have labeled 108K sentences, 41% of which as having dialectal content. We also present experimental results on the task of automatic dialect identification, using the collected labels for training and evaluation.
- Added by :
- Notes:
- Lenguaje dominante
- HTML
- Estrellas
- 91
- Forks
- 47
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de bigscience-workshop/data_tooling
-
data catalog
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
-
data catalog
Dificultad 1/5 Menos de una hora Aptitud para principiantes 62/100
-
Create dataset bhaskar Abiertodata catalog
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
Create dataset mediapart Abiertodata catalog
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
-
Create dataset goierna_magazine Abiertodata catalog
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100