`tft.compute_and_apply_vocabulary` is not robust to `RaggedTensor` type
@iindyk 已经在做这个了。
开始于 2022年4月1日。
评估
这个 Issue 还没有评估数据。
描述
Tensorflow version 2.8.0, TFT version 1.7.0.
I am currently working on constructing a module which has some multivalent inputs, as well as a multi-hot label endpoint. Both of these need similar transform and feature engineering: convert a string into tokens then map the tokens into a sequence of integers which are fed to an embedding table. However, the number of tokens in a given example string is not constant, and tft.compute_and_apply_vocabulary seems to be unable to parse the output of tf.string.split. In the context of the full model:
def _preprocess_multivalent_feature(feature, ncats):
raw_values = tf.strings.split(feature, ',')
coded_values = tft.compute_and_apply_vocabulary(raw_values, num_oov_buckets=1)
return coded_values
which lands me at (snipped for brevity)
TypeError Traceback (most recent call last)
/opt/conda/lib/python3.7/site-packages/tensorflow/python/framework/tensor_util.py in make_tensor_proto(values, dtype, shape, verify_shape, allow_broadcast)
548 try:
--> 549 str_values = [compat.as_bytes(x) for x in proto_values]
550 except TypeError:
/opt/conda/lib/python3.7/site-packages/tensorflow/python/framework/tensor_util.py in <listcomp>(.0)
548 try:
--> 549 str_values = [compat.as_bytes(x) for x in proto_values]
550 except TypeError:
/opt/conda/lib/python3.7/site-packages/tensorflow/python/util/compat.py in as_bytes(bytes_or_text, encoding)
86 raise TypeError('Expected binary or unicode string, got %r' %
---> 87 (bytes_or_text,))
88
TypeError: Expected binary or unicode string, got tf.RaggedTensor(values=tf.RaggedTensor(values=Tensor("StringSplit/StringSplit/StringSplit/StringSplitV2:1", shape=(None,), dtype=string), row_splits=Tensor("StringSplit/StringSplit/StringSplit/RaggedFromValueRowIds/RowPartitionFromValueRowIds/concat:0", shape=(None,), dtype=int64)), row_splits=Tensor("StringSplit/RaggedFromTensor/RaggedFromUniformRowLength/RowPartitionFromUniformRowLength/mul:0", shape=(None,), dtype=int64))
During handling of the above exception, another exception occurred:
[snip]
/app/pipeline/components/transform.py in _preprocess_multivalent_feature(feature, ncats)
69 def _preprocess_multivalent_feature(feature):
70 raw_values = tf.strings.split(feature, ',')
---> 71 coded_values = tft.compute_and_apply_vocabulary(raw_values, num_oov_buckets=1)
72 return coded_values
[snip]
/opt/conda/lib/python3.7/site-packages/tensorflow/python/framework/tensor_util.py in make_tensor_proto(values, dtype, shape, verify_shape, allow_broadcast)
551 raise TypeError("Failed to convert object of type %s to Tensor. "
552 "Contents: %s. Consider casting elements to a "
--> 553 "supported type." % (type(values), values))
554 tensor_proto.string_val.extend(str_values)
555 return tensor_proto
TypeError: Failed to convert object of type <class 'tensorflow.python.ops.ragged.ragged_tensor.RaggedTensor'> to Tensor. Contents: tf.RaggedTensor(values=tf.RaggedTensor(values=Tensor("StringSplit/StringSplit/StringSplit/StringSplitV2:1", shape=(None,), dtype=string), row_splits=Tensor("StringSplit/StringSplit/StringSplit/RaggedFromValueRowIds/RowPartitionFromValueRowIds/concat:0", shape=(None,), dtype=int64)), row_splits=Tensor("StringSplit/RaggedFromTensor/RaggedFromUniformRowLength/RowPartitionFromUniformRowLength/mul:0", shape=(None,), dtype=int64)). Consider casting elements to a supported type.
Commenting out the tf.string.split (and thus leaving them as whole strings) allows the pipeline execution to continue. Despite efforts, I cannot reproduce this exactly with a smaller example (this is work to scale up the pipeline through TFX, and providing that is out of the scope of the issue here). I am able to produce a RaggedTensor with the output of a similar function in a working example with faked inputs. However, I have a hard time believing that the tensor produced by that example would be useable by the Embedding layer which it is putatively going to be connected to:
Raw data:
[{'x': ['a,b', 'b,c,d', '']}, {'x': ['a,b,c', 'd', 'e,f,g,h,i']}]
Transformed data:
[{'_x$ragged_values': array([3, 0, 0, 2, 1, 9]),
'_x$row_lengths_1': array([2, 3, 1])},
{'_x$ragged_values': array([3, 0, 2, 1, 8, 7, 6, 5, 4]),
'_x$row_lengths_1': array([3, 1, 5])}]
It is highly undesirable, though perhaps acceptable, if there is a way to generate a normal tensor from the ragged one. The change:
return coded_values -> return coded_values.to_tensor()
However, that meets the same problem, as it appears earlier in compute_and_apply_vocabulary.
Any advice is appreciated.
- 主要语言
- Python
- 星标
- 989
- 派生
- 225
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
tensorflow/transform 的其他 Issue
-
stat:contributions welcome type:bug
难度 2/5 1-3 小时 新手友好度 45/100
tensorflow/transform#347 ·
-
bug stat:awaiting response
难度 3/5 1-2 天 新手友好度 35/100
tensorflow/transform#339 · 2 条评论 ·
-
examples/README.md has dead link to getting started可能重新可做 @pindinagesh 于 1598 天前认领,目前没有进行中的 PR。 未关闭stat:contributions welcome type:support
tensorflow/transform#272 · 4 条评论 · 已指派 1 人 ·
-
scale_to_z_score_per_key should give caller control over OOV behavior可能重新可做 @iindyk 于 1791 天前认领,目前没有进行中的 PR。 未关闭stat:contributions welcome type:feature
tensorflow/transform#252 · 6 条评论 · 已指派 2 人 ·
-
Table not initialized when serving model可能重新可做 @varshaan 于 1985 天前认领,目前没有进行中的 PR。 未关闭Etsy stat:contributions welcome
tensorflow/transform#237 · 11 条评论 · 已指派 1 人 ·
查看 tensorflow/transform 的全部 Issue
相似的 Issue
-
enhancement P2
难度 2/5 1-3 小时 新手友好度 78/100
Toloka/tolokaforge#1776 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
TencentCloud/Octop#1622 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 68/100
-
Rust: `const _` gets its file's node ID, so the file node is relabelled `_` and gains a self-loop未关闭
难度 2/5 1-3 小时 新手友好度 78/100
Graphify-Labs/graphify#4064 · 1 条评论 ·
维护者通常 2 天内回复