Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Add dataset: royal_society_corpus

Aperta
#41 5 commenti 0 reazioni 1 assegnatario Vedi su GitHub

@shamikbose ci sta già lavorando.

Dal 12/7/2022.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

dataset
A URL for this dataset

https://fedora.clarin-d.uni-saarland.de/rsc/

Dataset description

The Royal Society Corpus (RSC) is based on the first two centuries of the Philosophical Transactions of the Royal Society of London from its beginning in 1665 to 1869. It includes all publications of the journal written mainly in English and containing running text. The Philosophical Transactions was the first periodical of scientific writing in England. Founded in 1665 by Henry Oldenburg, the first secretary of the Royal Society, it initially contained excerpts of letters of his scientific correspondence, reviews and summaries of recently-published books, and accounts of observations and experiments.

This offers an interesting dataset of text from the scientific domain across a long time period (1665-1869). Additionaly the dataset contains a range of annotations:

The corpus is tokenized and linguistically annotated for lemma and part-of-speech using TreeTagger (Schmid 1994, Schmid 1995). For spelling normalization we use a trained model of VARD (Baron and Rayson 2008). As a special feature, we encode with each unit (word token) its average surprisal, i.e. the average amount of information it encodes in number of bits, with words as units and trigram as contexts [cf. Genzel and Charniak 2002).
Detailed information on the linguistic and structural annotation of the RSC can be found here.

The RSC consists of approximately 35 million token and is encoded for text type (abstracts, articles), author, year of publication. Information about decade and 50-year periods are also available allowing for a diachronic analysis of different granularity. Token sizes of the different subcorpora and other corpus statistics can be found here.

Dataset modality

Text

Dataset licence

Creative Commons Attribution Non Commercial Share Alike 4.0 International

Other licence

No response

How can you access this data

As a download from a repository/website

Confirm the dataset has an open licence
  • To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian

[email protected]

Lingua principale
Nessun dato sulla lingua
Stelle
91
Fork
8
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di bigscience-workshop/lam

Tutte le issue di bigscience-workshop/lam

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.