DVB subtitle decoder silently discards character-coded ("string coding") object segments — coding_method == 1 is a no-op that reports success
Maintainers usually reply within 1 day
@GuTS805 is already working on this.
Since Aug 11, 2026.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- c
- Domain
- audio-video-rtc
Research direction
Start in src/lib_ccx/dvb_subtitle_decoder.c at dvbsub_parse_object_segment() and trace its callers, dvbsub_decode() and dvbsub_parse_segment. Reproduce the coding_method == 1 case using the described crafted segment buffers, then consult ETSI EN 300 743 §7.2.5.2. Done means character-coded objects are handled or their loss is surfaced distinctly from successful decoding, with coverage for the affected path.
Written by the indexing model from the issue text.
Description
Component
C-core — src/lib_ccx/dvb_subtitle_decoder.c
Problem
dvbsub_parse_object_segment() branches on object_coding_method (bits 3-2 of the
object-coding-method byte, per ETSI EN 300 743 §7.2.5). coding_method == 0 (pixel/bitmap
objects) is fully handled via dvbsub_parse_pixel_data_block(). But coding_method == 1
("character coded" / string-coded objects — a legal, bandwidth-saving encoding some
broadcasters use) is a no-op:
else if (coding_method == 1)
{
mprint("FIXME support for string coding standard\n");
}
else
{
mprint("Unknown object coding %d\n", coding_method);
}
return 0;
No pixel or text data is ever written for that object, and the function still return 0
— the same value used on success. Every caller up the chain (dvbsub_decode() →
dvbsub_parse_segment switch) treats 0 as "proceed normally," so the caption region is
silently left blank with only a buried mprint line as evidence.
Why it matters
CCExtractor's entire purpose is making caption/subtitle content accessible. This isn't a
crash or loud error — it's silent data loss. A hard-of-hearing/deaf viewer relying on
extracted captions gets nothing for that segment, and there's no clear signal in normal
usage (the FIXME line is easy to miss in a full extraction log, and nothing marks the
output as incomplete).
Reproduction (confirmed against current master, actual code — not simulated)
I built current master with autotools (./autogen.sh && ./configure && make) and wrote
a small harness that #includes dvb_subtitle_decoder.c directly (so it calls the real,
compiled static functions) and drives it with two crafted segment buffers per EN 300 743:
A minimal region-composition-segment (16 bytes) registering region_id=1 (16×16, depth=4) and referencing object_id=1 — matches what a real broadcast stream sends before an object segment.
An object-data-segment (00 01 04) for object_id=1 with object_coding_method=1 (byte 0x04 → bits [3:2] = 01).
Output:
Region 1 registered OK. region->dirty BEFORE object segment = 0
FIXME support for string coding standard
dvbsub_parse_object_segment(coding_method=1) returned = 0 (0 == "success" to every caller)
region->dirty AFTER object segment = 0 (0 == no pixel/text data was ever written)
CONFIRMED: character-coded DVB subtitle object was silently discarded.
region->dirty (the flag dvbsub_parse_pixel_data_block() sets when it writes pixel
data) stays 0 — confirming no output was produced — while the function still returns
success.
Not a duplicate
Searched open/closed issues and PRs for "string coding", "coding_method",
"dvb subtitle object", "character-coded", and dvb_subtitle_decoder. The closest
historical issue, #243 ("Corrupt or empty subtitles (OCR, ts, DVB)"), has unrelated root
causes (missing DVBSUB_DISPLAY_SEGMENT handling, OCR failures, a telxcc fuzzy-match
bug). The one PR touching this file, #1794, only added malloc NULL checks.
Proposed fix direction
Minimum: stop returning 0 (success) from a path that did no work — surface this at a real warning/error level distinguishable from "decoded successfully" (e.g. track dropped-object stats so users can tell captions were silently lost).
Full fix: implement §7.2.5.2 character-object parsing (map character codes through the associated CLUT/character table). Since CCExtractor already OCRs pixel-object bitmaps down to text, a character-coded object could arguably be extracted as plain text with less code than the existing bitmap path.
- Dominant language
- C
- Stars
- 901
- Forks
- 592
- Avg merge
- 5d 19h
- Merged PRs (30d)
- 5
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from CCExtractor/ccextractor
-
[BUG] Legacy options -608, -708, -90090, -UCLA, -CC2, -LF, -DF, -parsepat, -parsepmt still rejected after #1856Possibly taken @Deepak-negi11 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
CCExtractor/ccextractor#2367 ·
Maintainers usually reply within 1 day
-
[BUG] Memory leak in free_sub_track(): blockaddition and message buffer never freed for WebVTT tracksPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 1/5 Under an hour Newbie friendliness 86/100
CCExtractor/ccextractor#2247 ·
Maintainers usually reply within 1 day
-
`--out=mcc`: CDP cc_count field overflows above 31 triplets; uint8 `data_size` corrupts lengths and over-reads at higher countsPossibly taken @kaihere14 claimed this 8 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 72/100
CCExtractor/ccextractor#2365 · 1 comment ·
Maintainers usually reply within 1 day
-
[BUG] MAX_CC_COUNT = 31 in ccxr_process_cc_data silently discards CEA-708 data on H.264 frames with multiple SEI messagesPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 4/5 3-5 days Newbie friendliness 68/100
CCExtractor/ccextractor#2358 · 1 comment ·
Maintainers usually reply within 1 day
-
--tpages-all extracts less than --tpagePossibly taken @SajalDevX claimed this 16 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 65/100
CCExtractor/ccextractor#2355 ·
Maintainers usually reply within 1 day
All issues in CCExtractor/ccextractor
Similar issues
-
backlog
Difficulty 1/5 Under an hour Newbie friendliness 82/100
EchoTools/nevr-runtime#454 ·
Maintainers usually reply within 1 day
-
initramfs: -type f (#18686) skips the libcurl.so.4 symlink, libcurl no longer copied into initramfsOpen
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 2 days
-
common/json_parse: json_to_bitcoin_amount fails to detect overflow and accepts negative/empty inputsPossibly taken @bhuvan-somisetty claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
ElementsProject/lightning#9617 ·
Maintainers usually reply within 2 days
-
Difficulty 1/5 Under an hour Newbie friendliness 62/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
zephyrproject-rtos/zephyr#121795 ·
Maintainers usually reply within 2 days