Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Outputted files have erroneous characters [bug on mac os]

オープン
#17 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
python

調査の方向性

evaluator.py と、提供された train.py 評価コマンドで使用される create_reference_files パスから始めます。ids.java_sa-python_sa.text.txt と eval_scripts/java_sa-python_sa.test/ 配下のスクリプトがどのようにデコードされ、書き込まれるかを調べ、報告されたセットアップ上で提供されたコマンドを使って問題を再現します。生成されたテキストとスクリプトに誤った文字やスペースが含まれなくなり、computational accuracy の評価でそれらを解析できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

enhancement

Hi, I was trying to reproduce the results in the TransCoder paper, but ran into some issues when it computed the computational accuracy.
It seems that when writing things such as id’s or outputted programs to files or outputs in the log, the text has some issues.

For example, in ids.java_sa-python_sa.text.txt (created by create_reference_files in evaluator.py), the lines look like “CHECK_@@ WHE@@ THER_@@ GI@@ V@@ EN_@@ NUMBER_@@ EV@@ EN_@@ O@@ DD”.

The scripts outputted by the model (e.g. in eval_scripts/java_sa-python_sa.test/) that are used to compute the computational accuracy similarly have erroneous characters and spaces (which causes syntax errors), e.g.:

def f_filled is_@@ ap ( arr , n ) :
    if n == 1 : return True
    arr.sort ( )
    d = arr [ 1 ] - arr [ 0 ]
    for i in range ( 2 , n ) :
        if arr [ i ] - arr [ i - 1 ] != d : return False
    return True

If it is relevant, I am on Mac OS Catalina with python 3.9 and this is the command I have been running to evaluate the TransCoder models provided:

python codegen_sources/model/train.py \
--eval_only True \
--reload_model 'TransCoder_model_1.pth,TransCoder_model_2.pth' \
--data_path "test_dataset" \
--exp_name transcoder \
--dump_path 'dump' \
--lgs 'java_sa-python_sa'  \
--bt_steps 'python_sa-java_sa-python_sa,java_sa-python_sa-java_sa'  \
--ae_steps 'python_sa,java_sa'  \
--mt_steps 'java_sa-python_sa,python_sa-java_sa' \
--encoder_only False \
--emb_dim 1024 \
--n_heads 8 \
--n_layers 0 \
--n_layers_encoder 6  \
--n_layers_decoder 6 \
--eval_bleu true \
--eval_computation true \
--has_sentences_ids true

Thank you.

主要言語
Python
スター
777
フォーク
144
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

facebookresearch/CodeGen のほかの issue

facebookresearch/CodeGen の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。