Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[BUG] store_hdcc() silently overwrites buffered caption data instead of concatenating when the same seq_index is reused with an unchanged timestamp

Open
#2,350 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@kaihere14 is already working on this.

Since Sep 17, 2026.

  • #2351 by @shanyuduo — closed without merging

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
c, rust

Research direction

Start in src/lib_ccx/sequencing.c at store_hdcc(), then trace calls from src/rust/src/es/userdata.rs and extension_and_user_data() in eau.rs. Inspect process_cc_data and build or find a focused reproduction with repeated same-timestamp user-data blocks; done means buffered blocks are preserved and downstream decoding handles the resulting triplets without regressions.

Written by the indexing model from the issue text.

Description

Summary

store_hdcc() in src/lib_ccx/sequencing.c:23-83 is meant to buffer caption blocks extracted from video for later reordering (B-frames can arrive out of temporal order, so captions get parked by sequence index and flushed in the correct order by process_hdcc()). A comment in the function states the intent explicitly: "Changed by CFS to concat, i.e. don't assume there's no data already for this seq_index."

The code no longer does this. When store_hdcc() is called twice for the same seq_index with the same timestamp (cc_fts[seq_index] == current_fts_now, so the flush-before-overwrite check doesn't trigger), the second call silently erases the first call's buffered caption data instead of appending to it.

Root cause
if (dec_ctx->cc_data_count[seq_index] > 0)
{
    if (dec_ctx->has_ccdata_buffered && dec_ctx->cc_fts[seq_index] != current_fts_now)
    {
        process_hdcc(enc_ctx, dec_ctx, sub);   // only runs if fts differs — NOT this case
    }
}
dec_ctx->cc_fts[seq_index] = current_fts_now;
dec_ctx->cc_data_count[seq_index] = 0;          // reset happens unconditionally, before the write
memcpy(dec_ctx->cc_data_pkts[seq_index] + dec_ctx->cc_data_count[seq_index] * 3, cc_data, cc_count * 3 + 1);
// offset is always 0 * 3 = 0, since count was just zeroed above — always writes to the start of the buffer
dec_ctx->cc_data_count[seq_index] += cc_count;  // = 0 + new count, since count was reset — functionally a plain assignment, not an accumulation

cc_data_count[seq_index] is reset to 0 unconditionally, immediately before that same value is used to compute the memcpy write offset. The offset is therefore always 0, so any second write to the same slot overwrites from the start of the buffer rather than appending after existing data. The += cc_count on the following line is effectively dead code in this path, since the count was just zeroed — it always evaluates to a plain assignment of the new count, never a true accumulation.

This wasn't always the behavior — traced via git history

The original 2014 implementation (src/sequencing.cpp, commit 6e0db8aa) had this guard:

if (stream_mode != CCX_SM_MP4) // CFS: Very ugly hack, but looks like overwriting is needed for at least some ES
    cc_data_count[seq_index] = 0;
memcpy(cc_data_pkts[seq_index] + cc_data_count[seq_index] * 3, cc_data, cc_count * 3 + 1);

The reset was conditional: only for non-MP4 (elementary stream) sources. For MP4, the count was deliberately left alone, producing a real concat, the offset naturally landed after existing data. This matches the comment ("concat... needed at least for MP4 samples") exactly, and isn't contradictory in this original form.

The guard was disabled in commit 6aac9dad433a (Anshul Maheshwari, 2015-07-21, commit message: "not working multiprogram"). That commit changed store_hdcc's context parameter from lib_ccx_ctx* to lib_cc_decode* as part of a multiprogram refactor. lib_cc_decode had no demux_ctx, so stream_mode could no longer be computed, and rather than re-plumbing it through, the stream_mode computation and the guard around the reset were both commented out, leaving the reset unconditional:

-    stream_mode = ctx->demux_ctx->get_stream_mode(ctx->demux_ctx);
+    //stream_mode = ctx->demux_ctx->get_stream_mode(ctx->demux_ctx);
...
-                    if (stream_mode!=CCX_SM_MP4) // CFS: Very ugly hack, but looks like overwriting is needed for at least some ES
+                    //if (stream_mode!=CCX_SM_MP4) // CFS: Very ugly hack, but looks like overwriting is needed for at least some ES
                             cc_data_count[seq_index]  = 0;

Every later commit touching these lines since (2019, 2020, 2022) is a rename or a clang-format pass, none touch this logic. This has been unconditional for roughly 11 years.

I don't think this was a deliberate decision to change the concat behavior, the commit's own message describes it as not working / incomplete, and it touches only the reset guard as a side effect of an unrelated struct-type change, not a considered removal of concat support. I want to be careful not to overstate that though, I can't verify intent with certainty, only what the commit history shows.

stream_mode/CCX_SM_MP4 no longer exists as a field on lib_cc_decode today (grepped lib_ccx.h, zero hits), so restoring the original guard verbatim isn't directly possible without re-plumbing that information back in, or finding a different way to distinguish the cases that guard was protecting.

Reachability, this isn't theoretical

The live code path today is Rust's user_data() in src/rust/src/es/userdata.rs, called from extension_and_user_data() in eau.rs (the C equivalent, src/lib_ccx/es_userdata.c's user_data(), is dead/unreferenced, confirmed by grep). extension_and_user_data() loops over every consecutive 0xB2 user-data start code within a single picture, calling user_data() once per occurrence. Multiple caption-producing formats can coexist and are each handled in this same dispatcher (GA94/HDTV, SCTE-20, Dish headers), and each ends in its own store_hdcc(...) call using the same picture-scoped current_tref/fts_now, the same seq_index and fts.

Any single picture carrying two such caption-producing user_data blocks (duplicate/redundant GA94 insertion, or GA94 combined with SCTE-20 or a Dish header in the same picture) hits exactly this same-seq_index/same-fts collision. This is a structurally reachable path in the currently-shipping decoder.

Caveat: I've confirmed this mechanism by tracing the source, but I have not yet caught it against a real captured file showing visibly missing captions. I'd like to try to build or find a repro, but wanted to file this now with the mechanism and root cause fully evidenced rather than wait.

Open question I haven't resolved

If store_hdcc is fixed to properly concat, I don't yet know whether the decoder side (process_cc_data and downstream) correctly handles a buffer containing two concatenated caption envelopes worth of triplets in one slot, versus expecting exactly one envelope's worth. process_cc_data iterates through cc_data_count[seq] triplets in a simple loop, which may just work regardless of how many logical envelopes contributed the bytes, since 608/708 data is fundamentally a stream of 3-byte units, but I haven't confirmed this. If the decoder can't cleanly handle that, fixing the writer alone might just move where the data gets mishandled rather than actually fixing it.

Questions for you
  1. Does this match your memory of why the MP4-vs-ES distinction existed in the original guard? Given stream_mode no longer exists on this struct, would the right fix be re-plumbing that information back through, or is there a simpler condition that captures what that guard was actually protecting against?
  2. Should I first confirm whether the decoder side correctly handles concatenated multi-envelope data before proposing a fix to the writer side, or are these independent enough to fix separately?

Happy to keep digging on either the empirical repro or the decoder-side question, whichever would be more useful.

Dominant language
C
Stars
901
Forks
592
Avg merge
5d 19h
Merged PRs (30d)
5

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from CCExtractor/ccextractor

All issues in CCExtractor/ccextractor

Similar issues

More C issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.