Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Create dataset TALAA

Abierto
#290 4 comentarios 0 reacciones 1 asignado Ver en GitHub

@apergo-ai ya está trabajando en esto.

Desde el 10/1/2022.

Evaluación

Este issue todavía no se ha evaluado.

Descripción

data catalog need custodian permission

Source: Masader Project

  • uid: talaa
  • entry: https://arbml.github.io/masader/card.html?54
  • Link: https://github.com/saidziani/Arabic-News-Article-Classification
  • License : unknown
  • Year: 2015
  • Language: ar
  • Dialect: ar-MSA: (Arabic (Modern Standard Arabic))
  • Domain: news articles
  • Form: text
  • Collection Style: crawling
  • Description: collections of articles
  • Volume: 57,827
  • Unit: documents
  • Ethical Risks: Low
  • Provider: USTHB Algeria
  • Derived From:
  • Paper Title: Building TALAA, a Free General and Categorized Arabic Corpus
  • Paper Link: https://www.scitepress.org/Papers/2015/53521/53521.pdf
  • Script: Arab
  • Tokenized: No
  • Host: GitHub
  • Access: Free
  • Cost:
  • Test Split: Yes
  • Tasks: Topic Classification
  • Evaluation Set?:
  • Venue Title: ICAART
  • Citations: 4
  • Venue Type: conference
  • Venue Name: Conference on Agents and Artificial Intelligence
  • authors: Essma Selab,A. Guessoum
  • affiliations: ,
  • abstract: Arabic natural language processing (ANLP) has gained increasing interest over the last decade. However,
    the development of ANLP tools depends on the availability of large corpora. It turns out unfortunately that
    the scientific community has a deficit in large and varied Arabic corpora, especially ones that are freely
    accessible. With the Internet continuing its exponential growth, Arabic Internet content has also been
    following the trend, yielding large amounts of textual data available through different Arabic websites. This
    paper describes the TALAA corpus, a voluminous general Arabic corpus, built from daily Arabic
    newspaper websites. The corpus is a collection of more than 14 million words with 15,891,729 tokens
    contained in 57,827 different articles. A part of the TALAA corpus has been tagged to construct an
    annotated Arabic corpus of about 7000 tokens, the POS-tagger used containing a set of 58 detailed tags. The
    annotated corpus was manually checked by two human experts. The methodology used to construct TALAA
    is presented and various metrics are applied to it, showing the usefulness of the corpus. The corpus can be
    made available to the scientific community upon authorisation.
  • Added by :
  • Notes: The github and the paper states that they are 51K articles but the actual one is 83 articles
Lenguaje dominante
HTML
Estrellas
91
Forks
47
Métricas de merge de PR
Sin PR fusionados en 30 d

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de bigscience-workshop/data_tooling

Todos los issues de bigscience-workshop/data_tooling

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.