Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Create dataset AOC

Abierto
#287 4 comentarios 0 reacciones 1 asignado Ver en GitHub

@apergo-ai ya está trabajando en esto.

Desde el 10/1/2022.

Evaluación

Este issue todavía no se ha evaluado.

Descripción

data catalog need data sourcing feedback

Source: Masader Project

  • uid: arabic_online_commentary
  • entry: https://arbml.github.io/masader/card.html?39
  • Link: https://github.com/sjeblee/AOC
  • License : unknown
  • Year: 2011
  • Language: ar
  • Dialect: other
  • Domain: news articles
  • Form: text
  • Collection Style: crawling and annotation(other)
  • Description: a 52M-word monolingual dataset rich in dialectal content
  • Volume: 108,000
  • Unit: sentences
  • Ethical Risks: Low
  • Provider: Johns Hopkins University
  • Derived From:
  • Paper Title: The Arabic Online Commentary Dataset: an Annotated Dataset of Informal Arabic with High Dialectal Content
  • Paper Link: https://aclanthology.org/P11-2007.pdf
  • Script: Arab
  • Tokenized: No
  • Host: GitHub
  • Access: Free
  • Cost:
  • Test Split: No
  • Tasks: dialect identification
  • Evaluation Set?:
  • Venue Title: ACL
  • Citations: 147
  • Venue Type: conference
  • Venue Name: Assofications of computation linguisitcs
  • authors: Omar Zaidan,Chris Callison-Burch
  • affiliations: ,
  • abstract: The written form of Arabic, Modern Standard Arabic (MSA), differs quite a bit from the spoken dialects of Arabic, which are the true "native" languages of Arabic speakers used in daily life. However, due to MSA's prevalence in written form, almost all Arabic datasets have predominantly MSA content. We present the Arabic Online Commentary Dataset, a 52M-word monolingual dataset rich in dialectal content, and we describe our long-term annotation effort to identify the dialect level (and dialect itself) in each sentence of the dataset. So far, we have labeled 108K sentences, 41% of which as having dialectal content. We also present experimental results on the task of automatic dialect identification, using the collected labels for training and evaluation.
  • Added by :
  • Notes:
Lenguaje dominante
HTML
Estrellas
91
Forks
47
Métricas de merge de PR
Sin PR fusionados en 30 d

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de bigscience-workshop/data_tooling

Todos los issues de bigscience-workshop/data_tooling

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.