Silent failure with --runner=DataflowRunner
还没有人认领这个 Issue。
评估
这个 Issue 还没有评估数据。
描述
System information
- Environment: Google Colab, Vertex AI (
KubeflowV2DagRunner) - TensorFlow version: 2.8.0
- TFX Version: 1.7.0
- Python version: 3.7.12
Here's my preprocessing_fn (redacted for clarity):
_FEATURES = [# list of str
]
_SPECIAL_IMPUTE = {
'special_foo': 1,
}
HOURS = [1, 2, 3, 4]
TABLE_KEYS = {
'XXX': ['XXX_1', 'XXX_2', 'XXX_3'],
'YYY': ['YYY_1', 'YYY_2', 'YYY_3'],
}
@tf.function
def _divide(a, b):
return tf.math.divide_no_nan(tf.cast(a, tf.float32), tf.cast(b, tf.float32))
def preprocessing_fn(inputs):
x = {}
for name, tensor in sorted(inputs.items()):
if tensor.dtype == tf.bool:
tensor = tf.cast(tensor, tf.int64)
if isinstance(tensor, tf.sparse.SparseTensor):
default_value = '' if tensor.dtype == tf.string else 0
tensor = tft.sparse_tensor_to_dense_with_shape(tensor, [None, 1], default_value)
x[name] = tensor
x['foo'] = _divide((x['foo1'] - x['foo2']), x['foo_denom'])
x['bar'] = tf.cast(x['bar'] > 0, tf.int64)
for hour in HOURS:
total = tf.constant(0, dtype=tf.int64)
for device_type in DEVICE_TYPES.keys():
total = total + x[f'some_device_{device_type}_{hour}h']
# one hot encode categorical values
for name, keys in TABLE_KEYS.items():
with tf.init_scope():
initializer = tf.lookup.KeyValueTensorInitializer(
tf.constant(keys),
tf.constant([i for i in range(len(keys))]))
table = tf.lookup.StaticHashTable(initializer, default_value=-1)
indices = table.lookup(tf.squeeze(x[name], axis=1))
one_hot = tf.one_hot(indices, len(keys), dtype=tf.int64)
for i, _tensor in enumerate(tf.split(one_hot, num_or_size_splits=len(keys), axis=1)):
x[f'{name}_{keys[i]}'] = _tensor
return {name: tft.scale_to_0_1(x[name]) for name in _FEATURES}
Here's the beam_pipeline_args:
BIG_QUERY_WITH_DIRECT_RUNNER_BEAM_PIPELINE_ARGS = [
'--project=' + GOOGLE_CLOUD_PROJECT,
'--temp_location=' + os.path.join('gs://', GCS_BUCKET_NAME, 'tmp'),
'--runner=DataflowRunner',
'--region=us-central1',
'--experiments=upload_graph', # must be enabled, otherwise fails with 413
'--dataflow_service_options=enable_prime',
'--autoscaling_algorithm=THROUGHPUT_BASED',
]
Not sure if related but with the above preprocessing_fn, my transform first failed with the error:
RuntimeError: The order of analyzers in your `preprocessing_fn` appears to be non-deterministic. This can be fixed either by changing your `preprocessing_fn` such that tf.Transform analyzers are encountered in a deterministic order or by passing a unique name to each analyzer API call.
I then added names to the tft.scale_to_0_1 analyzers:
return {name: tft.scale_to_0_1(x[name], name=f'{name}_scale_to_0_1') for name in _FEATURES}
After which my transform just silently failed without logs (see first screenshot). I check the worker logs but there's nothing substantial, only warnings (see second screenshot).
It's worth noting that I have the enable_prime flag.
- 主要语言
- Python
- 星标
- 989
- 派生
- 225
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
tensorflow/transform 的其他 Issue
-
stat:contributions welcome type:bug
难度 2/5 1-3 小时 新手友好度 45/100
tensorflow/transform#347 ·
-
bug stat:awaiting response
难度 3/5 1-2 天 新手友好度 35/100
tensorflow/transform#339 · 2 条评论 ·
-
examples/README.md has dead link to getting started可能重新可做 @pindinagesh 于 1599 天前认领,目前没有进行中的 PR。 未关闭stat:contributions welcome type:support
tensorflow/transform#272 · 4 条评论 · 已指派 1 人 ·
-
scale_to_z_score_per_key should give caller control over OOV behavior可能重新可做 @iindyk 于 1791 天前认领,目前没有进行中的 PR。 未关闭stat:contributions welcome type:feature
tensorflow/transform#252 · 6 条评论 · 已指派 2 人 ·
-
Table not initialized when serving model可能重新可做 @varshaan 于 1986 天前认领,目前没有进行中的 PR。 未关闭Etsy stat:contributions welcome
tensorflow/transform#237 · 11 条评论 · 已指派 1 人 ·
查看 tensorflow/transform 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 91/100
-
难度 1/5 1-3 小时 新手友好度 92/100
-
enhancement P2
难度 2/5 1-3 小时 新手友好度 78/100
Toloka/tolokaforge#1776 ·
维护者通常 1 天内回复
-
[bug] 本地会话日志回退在事件循环上同步读取 sessions/*.jsonl,长会话下阻塞同 loop 请求可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
TencentCloud/Octop#1622 ·
维护者通常 1 天内回复
-
arch area:fleet priority:p3 severity:low track:hosted-product
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复