Add dataset: french_fiction_16_18th_century
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 68/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- data
Research direction
Start by reviewing the repository's existing dataset contribution structure, then use the Zenodo URL and metadata in issue #86 as the source for the new entry. Done means the French fiction dataset is added with its description, text modality, licence, access method, and size recorded consistently with other datasets.
Written by the indexing model from the issue text.
Description
A URL for this dataset
https://zenodo.org/record/5770866
Dataset description
A corpus containing all digitized French novels from the beginning of print (the first entry is from 1473) to the 18th century.
French novels of the period have been identified using the Y2 quote of the French National Library Catalog that has served to classify past and present collections of novels in France from 1730 to 1996. Combined use of digitized sources from Gallica, Google Books, Archive.org and other digital library made it possible to attain a high representativeness: 78% of the novels of the 1450-1600 and 68% of the novels of the 1600-1700 have been retrieved.
Dataset modality
Text
Dataset licence
Creative Commons Attribution 4.0 International
Other licence
No response
How can you access this data
As a download from a repository/website
size of dataset
500MB-2GB
Confirm the dataset has an open licence
- To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian
No response
- Dominant language
- No language data
- Stars
- 91
- Forks
- 8
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from bigscience-workshop/lam
-
dataset
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
bigscience-workshop/lam#65 · 1 comment ·
-
candidate-dataset
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
bigscience-workshop/lam#92 · 1 comment ·
-
dataset
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
bigscience-workshop/lam#87 ·
-
dataset
Difficulty 4/5 3-5 days Newbie friendliness 42/100
bigscience-workshop/lam#85 · 1 reaction ·
-
Add dataset: TexBiG Opendataset
Difficulty 3/5 1-2 days Newbie friendliness 35/100
bigscience-workshop/lam#84 ·
All issues in bigscience-workshop/lam
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
open-compass/VLMEvalKit#1698 ·
-
triage:deciding
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
open-telemetry/otel-arrow#4123 · 1 reaction ·