Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Lazy-load TensorFlow for TFRecord-specific pipeline paths

オープン
#151 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

@hanzalaareeb がすでに取り組んでいます。

2026年8月20日 から。

  • #152 @hanzalaareeb による — オープン
  • #208 @venuachari777-bot による — オープン

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
68/100
issue の種類
リファクタリング
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python, tensorflow

調査の方向性

まず dpsynth.data_generation と dpsynth.pipeline_transformations.input_output から始め、次に tfrecord_descriptor の import path と、TFRecord 固有の I/O 分岐を調べます。提供されている Python の import チェックを実行して、TFRecord 以外のモジュールが TensorFlow を sys.modules に残さないことを確認し、その後パイプラインのテストを実行して、既存の 11 個のテストがパスすることを確認します。

索引モデルが issue の本文から書いたものです。

説明

Summary

This PR removes eager TensorFlow initialization from shared pipeline import paths.

Previously, importing non-TFRecord pipeline functionality could transitively import TensorFlow:

dpsynth.data_generation
→ creating_data_recorder_converter
→ tfrecord_descriptor
→ tensorflow

As a result, workflows that do not use TFRecord, including the SWIFT pipeline test, still required TensorFlow to initialize successfully during module import.

This PR defers TFRecord-specific imports until DataFormat.TFRECORD is selected and imports TensorFlow only inside TFRecord-specific I/O paths.

Changes

  • Lazily import tfrecord_descriptor when DataFormat.TFRECORD is selected.
  • Move TensorFlow imports into TFRecord-specific branches in pipeline I/O.
  • Preserve TensorFlow type annotations using the existing from __future__ import annotations support and TYPE_CHECKING imports without requiring TensorFlow at module import time.
  • Add regression coverage to verify that importing non-TFRecord pipeline modules does not load TensorFlow.
  • Preserve existing public APIs and dependency extras.

Why

Python executes module-level imports when importing a module. The previous import structure therefore initialized TensorFlow during test collection:

pytest collection
→ import SWIFT test
→ import shared DPSynth modules
→ import TFRecord implementation
→ import TensorFlow

This happened before any SWIFT test executed.

The SWIFT workflow uses PipelineDP's LocalBackend and a dummy record converter and does not require TFRecord support. However, it could still be blocked if TensorFlow failed to initialize in the environment.

The change makes TensorFlow initialization conditional on actually entering a TFRecord code path:

shared pipeline import
→ select data format
   ├── non-TFRecord → TensorFlow not imported
   └── TFRECORD → load TFRecord implementation → import TensorFlow

This preserves TFRecord support while preventing unrelated pipeline workflows from depending on successful TensorFlow initialization.

Verification

Confirmed that importing the shared modules no longer loads TensorFlow:

python -c "import sys; \
from dpsynth import data_generation; \
from dpsynth.pipeline_transformations import input_output; \
assert 'tensorflow' not in sys.modules"

The previously blocked pipeline tests now complete successfully:

11 passed

Some unrelated JAX, Beam, httplib2, and Pyparsing warnings remain.

主要言語
Python
スター
32
フォーク
13
平均マージ
1日 19時間
マージ済み PR(30日)
20

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

google/dpsynth のほかの issue

google/dpsynth の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。