Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

add tidycjk and tidyEmoji

Aperta Adatta ai principianti
#15 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
72/100
Tipo di issue
Documentazione
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
r
Ambito
documentation

Direzione di ricerca

Inizia confrontando il modo in cui i pacchetti CRAN esistenti sono elencati nella vista delle attività Natural Language Processing, quindi esamina le pagine CRAN di tidycjk e tidyEmoji collegate nell'issue. Il lavoro è completato quando entrambi i pacchetti compaiono nella vista con descrizioni e link accurati e concisi.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Two packages that may fit this view, both tidy-style toolkits for text that whitespace-based tooling handles poorly.

tidycjk (0.1.0), for Chinese, Japanese and Korean. CJK writing does not separate words with whitespace, so a standard tokeniser returns either one undifferentiated blob per sentence or isolated characters. It provides cjk_segment() and cjk_tokens() with a pluggable segmenter backend (cjk_segmenters(), register_cjk_segmenter()), script and language identification (cjk_script(), cjk_blocks(), cjk_detect_language(), has_cjk(), cjk_ratio()), counts and summaries (cjk_char_counts(), cjk_summary()), and the width handling that matters whenever CJK text is aligned or truncated (cjk_width(), cjk_pad(), cjk_truncate(), to_fullwidth(), to_halfwidth()).

tidyEmoji (0.4.0), for emoji in text columns. Not every code point is an emoji and a single emoji can span several code points, so emoji statistics are fiddly to get right. It covers extraction (emoji_extract_unnest(), emoji_extract_nest(), emoji_tokens()), counting and density (emoji_frequency(), emoji_density(), emoji_ratio(), emoji_position()), categorisation (emoji_categorize()), sentiment and emotion scoring against published lexicons (emoji_sentiment(), emoji_emotion(), emoji_lexicons(), register_emoji_lexicon()), shortcode conversion both ways for preprocessing and accessibility (as_emoji_shortcode(), as_emoji_name(), emoji_to_text(), text_to_emoji(), emoji_sanitize()), context and collocation (emoji_context(), emoji_collocations(), emoji_ngrams(), emoji_cooccurrence()), and emoji_dfm() for document-by-emoji feature tables. emoji_token_cost() reports what emoji cost under a tokeniser, which matters for LLM preprocessing.

The sentiment lexicon is the Emoji Sentiment Ranking (Kralj Novak et al. 2015, doi:10.1371/journal.pone.0144296).

Disclosure: I am the author of both packages.

Lingua principale
Nessun dato sulla lingua
Stelle
6
Fork
6
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di cran-task-views/NaturalLanguageProcessing

Tutte le issue di cran-task-views/NaturalLanguageProcessing

Issue simili

Altre issue su Documentation

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.