[Bug Report] `remove_batch_dim=True` with batch `size > 1`: the three caching paths disagree
Maintainer thường phản hồi trong vòng 1 ngày
@Mudassiruddin7 đang làm issue này rồi.
Từ ngày 6/10/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Describe the bug
The docstrings say remove_batch_dim "only makes sense with batch_size=1 inputs", but only one of the three caching paths enforces that; the other two silently do different wrong things on batch > 1:
run_with_cache(..., remove_batch_dim=True)(ActivationCache path): raisesAssertionError: Cannot remove batch dimension from cache with batch size 2. This is loud and arguably the right behavior.run_with_cache(..., remove_batch_dim=True, return_cache_object=False): silently ignoresremove_batch_dim, activations keep their batch dim.get_caching_hooks(remove_batch_dim=True)/add_caching_hooks:save_hookdoesstored[0], silently discarding every example after the first. This is bad, as it quietly abandons data.
Code example
import torch
from transformer_lens.config import TransformerBridgeConfig
from transformer_lens.model_bridge import TransformerBridge
cfg = TransformerBridgeConfig(
d_model=32, d_head=16, n_heads=2, n_layers=1, n_ctx=8, d_vocab=16,
d_mlp=64, act_fn="gelu", normalization_type="LN", seed=0,
)
bridge = TransformerBridge.boot_native(cfg)
tokens = torch.randint(0, cfg.d_vocab, (2, 6))
name = "blocks.0.hook_out"
with torch.no_grad():
try:
bridge.run_with_cache(tokens, names_filter=name, remove_batch_dim=True)
except AssertionError as e:
print("ActivationCache:", e)
_, plain = bridge.run_with_cache(
tokens, names_filter=name, remove_batch_dim=True, return_cache_object=False
)
print("plain dict:", tuple(plain[name].shape))
cache, fwd_hooks, _ = bridge.get_caching_hooks(names_filter=name, remove_batch_dim=True)
with torch.no_grad(), bridge.hooks(fwd_hooks=fwd_hooks):
bridge.forward(tokens)
print("get_caching_hooks:", tuple(cache[name].shape))
Output on dev (1012730f):
ActivationCache: Cannot remove batch dimension from cache with batch size 2
plain dict: (2, 6, 32)
get_caching_hooks: (6, 32)
Expected: one behavior across all three – the ActivationCache assert is the obvious candidate (an assert tensor.size(0) == 1 in save_hook and in the return_cache_object=False squeeze). EDIT: We have made some slight adjustments to this expected behavior, see discussion below
System Info
- Source checkout installed with
uv sync,devat1012730f. - macOS (Apple Silicon), CPU only; Python 3.12.12; torch 2.11.0; transformers 5.13.0.
- No model downloads needed: reproduced with a random-init
TransformerBridge.boot_nativemodel.
Additional context
Found while verifying #1856. The silent [0] is inherited from the legacy HookedRootModule.get_caching_hooks, which has the same behavior. A fix should probably touch both.
Checklist
- I have checked that there is no similar issue in the repo (required)
- Ngôn ngữ chính
- Python
- Star
- 3.9k
- Fork
- 708
- Merge trung bình
- 1 ngày 17 giờ
- Pull request đã merge (30 ngày)
- 70
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của TransformerLensOrg/TransformerLens
-
[Proposal] RoBERTa masked-LM adapter for TransformerBridgeCó thể đã có người làm @Canonik đã nhận hôm nay. Đang mởcomplexity-moderate new-architecture TransformerBridge
TransformerLensOrg/TransformerLens#1870 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Proposal] Backward Lens: support gated MLP gate/ up/ down gradient factorsCó thể đã có người làm @janmenjayap đã nhận 7 ngày trước. Đang mởcomplexity-moderate enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1832 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Proposal] Sparse probing: optional groups argument so rows from one prompt can't straddle the splitCó thể đã có người làm @lorenzozanee đã nhận 13 ngày trước. Đang mởcomplexity-simple enhancement help wanted TransformerBridge
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
TransformerLensOrg/TransformerLens#1813 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug Report] _BLOCK_LIST_ATTRS hardcoded name list silently drops Raven's blocks from composition-score / head-label analysisCó thể đã có người làm @LightWork666 đã nhận 17 ngày trước. Đang mởbug complexity-moderate TransformerBridge
TransformerLensOrg/TransformerLens#1791 · 2 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Proposal] SVD Circuits: singular-vector decomposition of a head's QK/ OV into causally-validated subfunctionsCó thể đã có người làm @janmenjayap đã nhận 28 ngày trước. Đang mởcomplexity-high enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1767 · 3 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của TransformerLensOrg/TransformerLens
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Deepak3699/Ai_Mentor#244 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
btclib-org/btclib-node#1880 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
CONTRIBUTING.md: say how ticketless bug fixes and feature PRs are handledCó thể đã có người làm @khuisman đã nhận hôm nay. Đang mởv0.9.2
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
khuisman/mcp-gee-sweet#941 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
pyjanitor-devs/pyjanitor#1758 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug ready for review
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
odysseus-dev/odysseus#6641 ·
Maintainer thường phản hồi trong vòng 1 ngày