[FEATURE] Native HLS and DASH stream parsing and extraction
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Quiet
- Tech stack
- c
- Domain
- audio-video-rtc
Research direction
The issue does not identify implementation files, tests, or an entry point. Begin by locating the stream-input and subtitle-extraction paths, then define the manifest, chunk-processing, and live-stream behavior that would demonstrate completion.
Written by the indexing model from the issue text.
Description
Currently, extracting subtitles from HLS (.m3u8) or DASH (.mpd) streams requires a multi-step workaround where users must first download the entire video stream to disk using external tools, run CCExtractor on the massive local file, and then manually delete the video afterward.
Adding native support to parse HLS/DASH manifests and directly process network streams would drastically improve the user experience and efficiency in three major ways:
-
Zero Video Download for Separate Subtitle Tracks In most modern streaming environments, subtitles are provided as an entirely separate playlist in the manifest (e.g.,
#EXT-X-MEDIA:TYPE=SUBTITLESpointing to a WebVTT playlist). By parsing the master manifest, CCExtractor could identify these separate tracks and only download the subtitle chunks (kilobytes of data). It would completely ignore the heavy video and audio chunks, saving gigabytes of internet bandwidth and drastically speeding up extraction. -
Zero Storage Overhead for Embedded Subtitles For older streams or live TV broadcasts where captions (like CEA-608/708) are multiplexed directly inside the video chunks (MPEG-TS packets), CCExtractor would still need to download the video data. However, native support would allow CCExtractor to operate entirely in memory like downloading a chunk, extracting the captions, and immediately discarding the chunk. This eliminates the need for the user to have 10GB+ of free disk space just to extract a few kilobytes of text.
-
Real-Time Extraction for Live Streams If a broadcast is live (e.g., Twitch, live news), there is no "finished" file to download. Native HLS/DASH support would allow CCExtractor to "tune in" to the live stream, continuously poll the manifest for the latest chunks, and output the subtitles in real-time as the event happens.
- Dominant language
- C
- Stars
- 901
- Forks
- 592
- Avg merge
- 7d 23h
- Merged PRs (30d)
- 11
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from CCExtractor/ccextractor
-
[BUG] Legacy options -608, -708, -90090, -UCLA, -CC2, -LF, -DF, -parsepat, -parsepmt still rejected after #1856Possibly taken @Deepak-negi11 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
CCExtractor/ccextractor#2367 ·
Maintainers usually reply within 1 day
-
[BUG] Memory leak in free_sub_track(): blockaddition and message buffer never freed for WebVTT tracksPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 1/5 Under an hour Newbie friendliness 86/100
CCExtractor/ccextractor#2247 ·
Maintainers usually reply within 1 day
-
`--out=mcc`: CDP cc_count field overflows above 31 triplets; uint8 `data_size` corrupts lengths and over-reads at higher countsPossibly taken @kaihere14 claimed this 8 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 72/100
CCExtractor/ccextractor#2365 · 1 comment ·
Maintainers usually reply within 1 day
-
[BUG] MAX_CC_COUNT = 31 in ccxr_process_cc_data silently discards CEA-708 data on H.264 frames with multiple SEI messagesPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 4/5 3-5 days Newbie friendliness 68/100
CCExtractor/ccextractor#2358 · 1 comment ·
Maintainers usually reply within 1 day
-
--tpages-all extracts less than --tpagePossibly taken @SajalDevX claimed this 16 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 65/100
CCExtractor/ccextractor#2355 ·
Maintainers usually reply within 1 day
All issues in CCExtractor/ccextractor
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
bug good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
tmewett/BrogueCE#929 · 1 comment ·
Maintainers usually reply within 1 day
-
bug : find_key() compares kty against "ocy" instead of "oct", breaking kid-less HS256 verificationOpen
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
OpenPrinting/cups#1756 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
SteamGridDB/SGDBoop#147 ·
-
Doc: insert executor and incremental consolidation leave sorted runs, not globally sorted headsOpen
Difficulty 1/5 Under an hour Newbie friendliness 83/100
semantic-reasoning/wirelog#2133 ·
Maintainers usually reply within 1 day