`ModelMixin.from_pretrained` on a sharded checkpoint fails under `HF_HUB_OFFLINE=1` even when fully cached (`_get_checkpoint_shard_files` still calls `model_info`)
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 70/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
調査の方向性
src/diffusers/utils/hub_utils.py の _get_checkpoint_shard_files から始め、modeling_utils.py の from_pretrained が local_files_only をどのように渡すかを追ってください。loaders/lora_base.py のオフライン処理と比較し、キャッシュ済みのシャード分割チェックポイントと HF_HUB_OFFLINE=1 を使って issue の再現手順を実行してください。キャッシュからのオフライン読み込みが成功し、シャードが見つからない場合には既存のエラーが引き続き発生すれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
What happens. With HF_HUB_OFFLINE=1 set and the checkpoint fully cached, SomeModel.from_pretrained("<repo id>") raises huggingface_hub.errors.OfflineModeIsEnabled whenever the checkpoint is sharded (it has a *.safetensors.index.json). Single-file checkpoints load fine in the same setup. Passing local_files_only=True as well works around it. The env var alone is the documented way to run offline, and it is the only option when the call happens inside a library you do not control.
Why. _get_checkpoint_shard_files checks that the shards exist on the Hub before calling snapshot_download, and it skips that check only when local_files_only is truthy:
# src/diffusers/utils/hub_utils.py, main @ da1d382, lines 433-435
# If the repo doesn't have the required shards, error out early even before downloading anything.
if not local_files_only:
model_files_info = model_info(pretrained_model_name_or_path, revision=revision, token=token)
ModelMixin.from_pretrained passes the caller's local_files_only straight through (modeling_utils.py line 1297), and that is None when the user relies on HF_HUB_OFFLINE. So model_info runs, and huggingface_hub refuses the request because offline mode is on. Earlier steps in the same load already handle offline mode: revision resolution logs Could not reach the Hub ... Using cached commit hash, and the config and index come from the cache. Only this pre-check makes the load fail. The guard comes from #12005, which fixed the local_files_only=True case (#11948). The env-var case was reported in #11447, which was closed citing #11428, a pipeline PR that does not touch this function.
With huggingface_hub 1.33, snapshot_download called afterwards also uses the network when it gets an already-resolved revision and the cache has no trees/<commit>.json listing. That listing is missing from caches written by older huggingface_hub versions and from hand-copied caches. So the fix should also pass the offline state to snapshot_download, as shown below.
I can open a PR for this once a maintainer confirms the approach.
Reproduction
Loads a real sharded checkpoint, first online to fill a fresh cache, then offline with the same call:
import os, subprocess, sys, tempfile, textwrap
cache = tempfile.mkdtemp()
code = textwrap.dedent("""
from diffusers import FluxTransformer2DModel
m = FluxTransformer2DModel.from_pretrained("hf-internal-testing/tiny-flux-sharded", subfolder="transformer")
print("loaded", sum(p.numel() for p in m.parameters()), "params")
""")
env = dict(os.environ, HF_HOME=cache)
subprocess.run([sys.executable, "-c", code], env=env, check=True) # online: fills the cache
env["HF_HUB_OFFLINE"] = "1"
sys.exit(subprocess.run([sys.executable, "-c", code], env=env).returncode) # offline: same call fails
Output on main @ da1d382: the online call prints loaded 68484 params. The offline call fails with the trace under Logs. With the fix below, the offline call also prints loaded 68484 params.
Logs
loaded 68484 params
Could not reach the Hub (Cannot reach https://huggingface.co/api/models/hf-internal-testing/tiny-flux-sharded: offline mode is enabled. To disable it, please unset the `HF_HUB_OFFLINE` environment variable.). Using cached commit hash for 'hf-internal-testing/tiny-flux-sharded'.
Traceback (most recent call last):
File "<string>", line 3, in <module>
File ".../huggingface_hub/utils/_validators.py", line 89, in _inner_fn
return fn(*args, **kwargs)
File ".../diffusers/models/modeling_utils.py", line 1297, in from_pretrained
resolved_model_file, sharded_metadata = _get_checkpoint_shard_files(
File ".../diffusers/utils/hub_utils.py", line 435, in _get_checkpoint_shard_files
model_files_info = model_info(pretrained_model_name_or_path, revision=revision, token=token)
File ".../huggingface_hub/utils/_validators.py", line 89, in _inner_fn
return fn(*args, **kwargs)
File ".../huggingface_hub/hf_api.py", line 3322, in model_info
r = get_session().get(path, headers=headers, timeout=timeout, params=params)
[... httpx frames ...]
File ".../huggingface_hub/utils/_http.py", line 284, in hf_request_event_hook
raise OfflineModeIsEnabled(
huggingface_hub.errors.OfflineModeIsEnabled: Cannot reach https://huggingface.co/api/models/hf-internal-testing/tiny-flux-sharded/revision/main: offline mode is enabled. To disable it, please unset the `HF_HUB_OFFLINE` environment variable.
Suggested fix. In _get_checkpoint_shard_files, treat offline mode the same as local_files_only=True. This skips the model_info pre-check and makes snapshot_download read from the cache. A missing shard is still reported by the existing os.path.isfile check after the download step.
ignore_patterns = ["*.json", "*.md"]
# In offline mode the shards can only come from the cache, so behave as if `local_files_only=True`.
local_files_only = local_files_only or HF_HUB_OFFLINE
HF_HUB_OFFLINE is already imported in hub_utils.py, and loaders/lora_base.py line 298 uses the same local_files_only or HF_HUB_OFFLINE pattern. dynamic_modules_utils.py line 423 also calls model_info without an offline guard (custom code loaded from a Hub repo). I have not reproduced that path, so it is left out here.
System Info
- Diffusers version: 0.41.0.dev0 (main @ da1d382, 2026-10-05); latest release v0.40.0 has the same guard
- Platform: Linux-6.1 x86_64 (Amazon Linux 2023), glibc 2.34
- Python version: 3.11.14
- PyTorch version: 2.14.1+cpu (CPU only)
- Huggingface_hub version: 1.33.0
- Transformers version: 5.19.0.dev0
- Accelerate version: not installed
- Safetensors version: 0.8.0
- Using GPU in script?: No
- Using distributed or parallel set-up in script?: No
Who can help?
@sayakpaul @DN6
- 主要言語
- Python
- スター
- 34.6k
- フォーク
- 7.4k
- 平均マージ
- 4日 15時間
- マージ済み PR(30日)
- 49
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
huggingface/diffusers のほかの issue
-
UniPCMultistepScheduler fails in torch.linalg.solve under a float64 default dtype: the unit entry of rks takes the default dtype対応中かも @DawnofGenX が 5 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
huggingface/diffusers#14888 ·
メンテナーはふだん 1 日以内に返信
-
WanAnimatePipeline.get_i2v_mask() defaults device to "cuda", which raises on non-CUDA accelerators (NPU/XPU/MPS)対応中かも @li-lizhe が 8 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
huggingface/diffusers#14881 ·
メンテナーはふだん 1 日以内に返信
-
IndexError when preprocessing an empty image or video list対応中かも @MohammadHijjawi97 が 13 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
huggingface/diffusers#14864 ·
メンテナーはふだん 1 日以内に返信
-
Windows: check_ai.py fails decoding UTF-8 guides with the default locale対応中かも @tanvir-ux が 13 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
huggingface/diffusers#14837 ·
メンテナーはふだん 1 日以内に返信
-
TangentialClassifierFreeGuidance.is_conditional reads an attribute that does not exist対応中かも @dafahaha が 19 日前に担当しました。 オープンbug needs-env-info pipelines
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
huggingface/diffusers#14794 ·
メンテナーはふだん 1 日以内に返信
huggingface/diffusers の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
QuantEcon/lecture-python-programming#642 ·
メンテナーはふだん 1 日以内に返信
-
area/config area/profiles comp/cli needs-decision P3 sweeper:risk-compatibility type/feature
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NousResearch/hermes-agent#133697 ·
メンテナーはふだん 1 日以内に返信
-
enhancement needs-triage
難易度 1/5 1時間未満 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
core
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
vectorize-io/hindsight#5279 ·
メンテナーはふだん 1 日以内に返信