Create dataset malindomorph__morphological_dictionary_and_analyser_for_malay_indonesian
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- json
- Lĩnh vực
- data
Hướng nghiên cứu
Xác định mục dữ liệu danh mục được biểu diễn bởi malindomorph__morphological_dictionary_and_analyser_for_malay_indonesian.json và so sánh cấu trúc của mục đó với siêu dữ liệu được cung cấp. Xác nhận rằng bản ghi MALINDOMorph được thêm vào với nguồn, tính khả dụng, giấy phép, ngôn ngữ và các chi tiết về phương tiện được nêu, bao gồm cả các trường xác thực chưa được giải quyết.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
- uid: malindomorph__morphological_dictionary_and_analyser_for_malay_indonesian
- type: processed
- description:
- name: MALINDOMorph: Morphological dictionary and analyser for Malay/Indonesian
- description: Malay/Indonesian lacked an open wide-coverage dictionary that can be used for both NLP tasks and non-NLP purposes. The MALINDO Morph morphological dictionary is the first such dictionary. It provides morphological information (root, prefix, suffix, circumfix, reduplication) for roughly 232K surface forms. The entry forms are those found in the authoritative dictionaries in Malaysia (Kamus Dewan4) and Indonesia (Kamus Besar Bahasa Indonesia5) (core dictionary) as well as frequent words in the Leipzig Corpora Collection (Goldhahn et al., 2012) (expanded dictionary). The morphological analyses were checked by hand for all surface forms, except for (i) basic and di-forms in the expanded dictionary whose existence is predicted from the corresponding meN-active forms in the core dictionary and (ii) the case variants of the items in the core dictionary. This paper also discusses the morphological analyser that we developed to create our morphological dictionary. Our morphological analyser is more linguistically rigorous than previous morphological analysers and stemmers/lemmatizers such as MorphInd (Larasati et al., 2011) because it takes into account circumfixes, which have previously been neglected, largely due to a misunderstanding among NLP researchers that circumfixes are no more than combinations of a prefix and a suffix.
- homepage: chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/viewer.html?pdfurl=http%3A%2F%2Flrec-conf.org%2Fworkshops%2Flrec2018%2FW29%2Fpdf%2F8_W29.pdf&clen=201938&chunk=true
- validated: True
- languages:
- language_names:
- Indonesian
- language_comments:
- language_locations:
- Asia
- Indonesia
- validated: False
- language_names:
- custodian:
- name: Hiroki Nomoto
- in_catalogue:
- type: A university or research institution
- location: Japan
- contact_name: Hiroki Nomoto
- contact_email: nomoto@tufs.ac.jp
- contact_submitter: False
- additional: chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/viewer.html?pdfurl=http%3A%2F%2Flrec-conf.org%2Fworkshops%2Flrec2018%2FW29%2Fpdf%2F8_W29.pdf&clen=201938&chunk=true
- validated: False
- availability:
- procurement:
- for_download: No - but the current owners/custodians have contact information for data queries
- download_url:
- download_email:
- licensing:
- has_licenses: Yes
- license_text:
- license_properties:
- license_list:
- pii:
- has_pii: Yes
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- procurement:
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - their documentation/homepage/description is available
- primary_license: Unclear / I don't know
- primary_types:
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- other
- text_is_transcribed: No
- instance_type:
- instance_count:
- instance_size:
- validated: False
- category:
- fname: malindomorph__morphological_dictionary_and_analyser_for_malay_indonesian.json
- Ngôn ngữ chính
- HTML
- Star
- 91
- Fork
- 47
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của bigscience-workshop/data_tooling
-
data catalog
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 82/100
-
data catalog
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 62/100
-
Create dataset bhaskar Đang mởdata catalog
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
-
Create dataset mediapart Đang mởdata catalog
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
-
Create dataset goierna_magazine Đang mởdata catalog
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Tất cả issue của bigscience-workshop/data_tooling
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
TauricResearch/TradingAgents#1397 ·
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
TencentCloud/Octop#1085 · 1 bình luận ·
-
Sandy macrofaunal assemblages adjacent to rocky reefs on São Miguel Island (Azores, NE Atlantic) Đang mởdataset
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
iobis/obis-network-datasets#909 ·
-
Copy cohorts to keep old cohorts Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
OHDSI/CohortConstructor#774 ·