Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Optional — design suggestion, not a bug] Folder scan merges every `*_L*.bin` in a directory into one dataset

オープン
#29 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
matlab, python
領域
data

調査の方向性

Start with Python's PyraviewDataset._scan_folder and MATLAB's Dataset.scanFolder, then review the existing tests and docs/API.md mentioned in the issue. The maintainer must first choose between a prefix, a guard, or documentation-only approach; done means the decision is implemented consistently in both bindings or documented clearly.

索引モデルが issue の本文から書いたものです。

説明

enhancement optional question

This is not a bug report. The current behaviour is self-consistent, matches the MATLAB binding, and is correct under the convention of one dataset per folder. This is a suggestion about whether that convention should be enforced rather than assumed. Closing it as "working as intended" is a perfectly reasonable outcome.

Current behaviour

PyraviewDataset._scan_folder (and MATLAB's Dataset.scanFolder) selects level files with a wildcard prefix:

for full_path in sorted(glob.glob(os.path.join(self.folder_path, '*_L*.bin'))):
d = dir(fullfile(obj.FolderPath, '*_L*.bin'));

Nothing associates expA_L1.bin with expA_L2.bin as belonging to the same recording — the prefix is a wildcard, so every pyramid file in the directory is treated as a level of a single dataset.

What that looks like with two recordings in one folder

expA at 1000 Hz starting at 5 s, expB at 500 Hz starting at 900 s, each with two levels:

files on disk : ['expA_L1.bin', 'expA_L2.bin', 'expB_L1.bin', 'expB_L2.bin']

native_rate   : 1000.0   (expA is 1000, expB is 500)
start_time    : 5.0      (expA is 5.0, expB is 900.0)
files         : ['expA_L1.bin', 'expB_L1.bin', 'expA_L2.bin', 'expB_L2.bin']
decimations   : [10, 10, 100, 100]
rates         : [100.0, 50.0, 10.0, 5.0]
start times   : [5.0, 900.0, 5.0, 900.0]

get_data(5..15s, 50px) -> (0,)   | level chosen from: expB_L2.bin

Three consequences, all silent:

  1. The two recordings merge into one pyramid, with duplicate decimation factors (10, 10, 100, 100) sorted together as if they were levels of one dataset.
  2. native_rate, native_start_time, channels and data_type come from whichever file sorts first; the other recording's metadata is discarded.
  3. Level selection can cross datasets. Asking for expA's window selects expB_L2.bin, because 5 Hz is the coarsest rate meeting the demand — and expB starts at 900 s, so the read falls outside the file and returns empty.

Why this may well be fine as-is

  • The MATLAB binding behaves identically. Changing only Python would reintroduce exactly the kind of divergence #22 closed, so any change should land in both.
  • One dataset per folder is a reasonable convention, and may be the only layout in practice. The test suites assume it (each test writes one dataset into a fresh temp directory), which is realistic rather than an oversight.
  • The property-based constructor already sidesteps it. Passing files=[...] explicitly — the NDI use case — never consults the glob, so the ambiguity does not arise there.

Options, if it seems worth addressing

  1. An optional prefix argument. PyraviewDataset(folder, prefix='expA') → globs expA_L*.bin; MATLAB Dataset(folder, 'Prefix', "expA"). Backward compatible, makes intent explicit.
  2. A guard. Warn or raise when the scan finds more than one distinct filename prefix, or duplicate decimation factors. Cheap, and turns a silent wrong answer into a clear message. Could be done without option 1.
  3. Documentation only. State the one-dataset-per-folder assumption in docs/API.md and the docstrings, and leave the behaviour alone.

My suggestion would be 2 on its own, or 1+2 together — 2 is what converts the failure mode from "quietly returns the wrong data" into something a user can act on. But this is a judgement call about how the library is actually used, which is yours to make.

Noticed while verifying the v0.4.1 wheels: a scratch directory held pyramid files from two earlier checks, and get_data returned an empty array rather than the expected samples.

主要言語
MATLAB
スター
0
フォーク
0
平均マージ
30分
マージ済み PR(30日)
6

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

似ている issue

MATLAB の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。