Feature request: Modify `text.regex_split_with_offsets()` behavior to be in line with `tf.strings.length()`
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 42/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- cpp, tensorflow
- Domain
- api, machine-learning
Research direction
Start with the text.regex_split_with_offsets() API and compare its documented offsets with tf.strings.length() and tf.strings.substr(), especially their BYTE and UTF8_CHAR behavior. Trace the implementation and existing tests for the split operation; done means the offsets use the requested unit and return tf.int32 values consistently with the related TensorFlow string APIs.
Written by the indexing model from the issue text.
Description
text.regex_split_with_offsets() currently returns begin and end as tf.int64 tensors that count indices in bytes.
tf.strings.length() on the other hand, returns a tf.int32 tensor which counts lengths in either bytes or UTF8 characters according to the value of the parameter unit.
So this would actually be two separate requests:
- Change the return types of
text.regex_split_with_offsets()totf.int32, removing the need for a cast when comparing withtf.strings.length(). I doubt there will be a use case for strings longer than INT32_MAX in the foreseeable future. - Add parameter
unit: Literal["BYTE", "UTF8_CHAR"] = "BYTE"matching the behavior oftf.strings.length()andtf.strings.substr(). Seeing the regular expressions are already being interpreted in 'utf-8', I think it would make sense to add a layer of abstraction to facilitate slicing by UTF-8 character index.
- Dominant language
- C++
- Stars
- 1.3k
- Forks
- 378
- Avg merge
- 10h 48m
- Merged PRs (30d)
- 2
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from tensorflow/text
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
tensorflow/text#1498 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
tensorflow/text#1497 ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
tensorflow/text#1448 · 3 reactions ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
tensorflow/text#1421 · 2 comments · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 38/100
tensorflow/text#1393 · 1 reaction ·
Maintainers usually reply within 1 day
Similar issues
-
agent:WSL bug linux LOW ui
Difficulty 1/5 Under an hour Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Copter: PosHold brake-entry threshold became 16 deg instead of 0.16 deg after the radians conversionOpen
Difficulty 1/5 Under an hour Newbie friendliness 78/100
ArduPilot/ardupilot#34617 · 1 comment · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
tesseract-robotics/tesseract_nanobind#168 ·
Maintainers usually reply within 1 day
-
Self-hosted runner Dockerfile pins actions/runner 2.327.1, below GitHub's new minimum (2.329.0)Openauto-triaged bug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
microsoft/microsoft-ui-xaml#12158 ·
Maintainers usually reply within 1 day