Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

`middle-word-em` extra with `allowed=False` breaks bold and leaks emphasis

Đang mở
#679 2 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
58/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Đình trệ
Công nghệ
python
Lĩnh vực
tooling

Hướng nghiên cứu

Bắt đầu trong markdown2.py với lớp MiddleWordEm và tái hiện các trường hợp được báo cáo bằng script Python được cung cấp, đặc biệt là các dòng in đậm liền kề và dấu gạch dưới giữa từ ở gần phần nhấn mạnh độc lập. So sánh HTML được render với các đầu ra mong đợi và xác minh toàn bộ ma trận kiểm thử, bao gồm inline code và các tên dunder, trước khi coi issue là hoàn tất.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Bug

Bug Report: middle-word-em extra with allowed=False breaks bold and leaks emphasis

Package: markdown2
Version: 2.5.4
Severity: High — silently corrupts rendered output
Affects: Any consumer using middle-word-em with allowed=False (or equivalently extras={"middle-word-em": False})


Summary

When the middle-word-em extra is configured with allowed=False to disable mid-word underscore emphasis (GFM behaviour), two bugs are triggered that silently corrupt the rendered HTML:

  1. Bold (**text**) rendering is broken across adjacent lines
  2. Unpaired mid-word underscores leak and create spurious <em> spans

Both bugs are in the MiddleWordEm class in markdown2.py.


Bug 1: Bold rendering broken on adjacent lines

Markdown Input
**Name**: Alice
**Role**: Engineer
**Department**: quality_assurance
Expected Output
<p><strong>Name</strong>: Alice
<strong>Role</strong>: Engineer
<strong>Department</strong>: quality_assurance</p>
Actual Output
<p><em>*Name<em>*: Alice
*</em>Role<em>*: Engineer
*</em>Department*</em>: quality_assurance</p>

Bold is entirely destroyed; the ** tokens are consumed by erroneous <em> matching.

Root Cause

When allowed=False, MiddleWordEm.__init__() wraps the inherited em_re with word-boundary assertions:

self.em_re = re.compile(r'(?<=\b)%s(?=\b)' % self.em_re.pattern, self.em_re.flags)

The original em_re is (\*|_)(?=\S)(.*?\S)\1. After wrapping it becomes:

(?<=\b)(\*|_)(?=\S)(.*?\S)\1(?=\b)

The problem: * is a non-word character (\W), so the transition from a word character to * constitutes a word boundary (\b). Consider two adjacent bold tokens:

**Name**: Alice\n**Role**: Engineer
  • The trailing * of **Name** follows e (word → non-word = \b) ✓ opens em
  • The leading * of **Role** precedes R (non-word → word = \b) ✓ closes em

So the wrapped em_re matches the entire span *: Alice\n* as emphasis, eating both bold delimiters and producing broken output.

This happens in ItalicAndBoldProcessor.run() at:

text = self.em_re.sub(self.sub, text)

And MiddleWordEm.sub() converts the match to <em>...</em>:

def sub(self, match):
    syntax = match.group(1)
    if len(syntax) != 1:
        return super().sub(match)     # strong → no-op
    return '<em>%s</em>' % match.group(2)  # em → wraps in <em>

Bug 2: Unpaired mid-word underscores leak into emphasis

Markdown Input
The function load_user_profile retrieves data.
Use _caution_ when calling it directly.
Expected Output
<p>The function load_user_profile retrieves data.
Use <em>caution</em> when calling it directly.</p>
Actual Output (even with Bug 1 worked around via a sub() no-op patch)
<p>The function load_user_profile<em>profile retrieves data.
Use _caution</em> when calling it directly.</p>

The _profile underscore inside the identifier is matched with the _caution_ emphasis delimiter, creating a spurious <em> span across both lines.

Root Cause

After the ItalicAndBoldProcessor.run() call, MiddleWordEm.run() hashes mid-word underscores using liberal_em_re:

text = self.liberal_em_re.sub(self.sub_hash, text)

The liberal_em_re pattern matches pairs of mid-word em characters:

self.liberal_em_re = re.compile(r'''
    (                # \1 - single em char in middle of word
        (?<![*_\s])  # not preceded by em char or whitespace
        [*_]         # em character
        (?![*_])     # not followed by another em char
    )
    (?=\S)           # must be followed by non-whitespace
    (.*?\S)          # content
    \1               # matching closing char
    (?!\s|$)         # must not be followed by whitespace or EOF
''', re.S | re.X)

For load_user_profile:

  • _user_ is matched and hashed (opening _ after d, closing _ before p) ✓
  • After hashing _user_, the text becomes load{hash}profile
  • The _ between {hash} and profile has no matching closing _ within the same word
  • So liberal_em_re does not hash it

Later, _do_italics_and_bold runs the standard em_re ((\*|_)(?=\S)(.*?\S)\1) which matches from the unprotected _profile to the standalone _caution_, creating the broken span.

The fundamental issue: liberal_em_re only protects matched pairs of mid-word underscores. Identifiers with an odd number of underscores (e.g., load_user_profile has 2 underscores; after pairing _user_, one _ is left over) leave unpaired underscores exposed to normal emphasis processing.


Proposed Fix

Replace MiddleWordEm.run() with a simpler approach that hashes every individual mid-word underscore rather than only matched pairs.

Current (buggy) run():
def run(self, text):
    if self.options['allowed']:
        return text

    # Bug 1: super().run() calls em_re.sub(self.sub, text) with the
    #         word-boundary-wrapped em_re, breaking **bold** rendering.
    text = super().run(text)

    if self.md.order < self.md.stage:
        # Bug 2: liberal_em_re only hashes matched PAIRS, leaving
        #         unpaired mid-word underscores unprotected.
        text = self.liberal_em_re.sub(self.sub_hash, text)
    return text
Proposed fix for run():
import re
_midword_underscore_re = re.compile(r'(?<=[a-zA-Z0-9])_(?=[a-zA-Z0-9])')

def run(self, text):
    if self.options['allowed']:
        return text

    if self.md.order < self.md.stage:
        # Hash every individual mid-word underscore (between alphanumeric
        # characters). This replaces both the buggy em_re.sub() call and
        # the incomplete liberal_em_re pair-matching approach.
        def _hash_midword(match):
            substr = match.group(0)
            key = _hash_text(substr)
            self.hash_table[key] = substr
            return key
        text = _midword_underscore_re.sub(_hash_midword, text)
    else:
        # After ITALIC_AND_BOLD stage: restore hashed underscores
        orig_text = ''
        while orig_text != text:
            orig_text = text
            for key, substr in self.hash_table.items():
                text = text.replace(key, substr)
    return text
Why this works

The regex (?<=[a-zA-Z0-9])_(?=[a-zA-Z0-9]) matches any underscore that has an alphanumeric character on both sides. This:

  1. Fixes Bug 1: Skips super().run() entirely (no em_re.sub() call), so the word-boundary-wrapped regex never runs against bold tokens.

  2. Fixes Bug 2: Hashes every individual mid-word underscore, not just matched pairs. For load_user_profile, all underscores are hashed independently, so none can participate in later emphasis matching.

  3. Preserves standalone emphasis: _caution_ has its opening _ preceded by a space (not alphanumeric), so it is NOT matched and remains available for normal emphasis processing.

  4. Preserves bold: **bold** contains no underscores and is untouched by the regex. The _do_italics_and_bold stage handles bold normally.

  5. Preserves __dunder__ syntax: Double underscores like __init__ have _ preceded by _ (not alphanumeric), so they are not matched.

  6. Safe inside code spans: By the time MiddleWordEm.run() executes (after CodeFriendly), inline code has already been hashed by markdown2's earlier processing stages. Underscores inside backtick code are not visible to this regex.


Test Matrix

# Input Expected Bug 1 (original) Bug 2 (sub no-op) Fixed
1 **Name**: Alice
**Role**: Engineer
<strong>Name</strong>: Alice
<strong>Role</strong>: Engineer
FAIL — bold destroyed PASS PASS
2 load_user_profile load_user_profile (literal) PASS PASS PASS
3 load_user_profile + _emphasis_ in same paragraph load_user_profile (literal) + <em>emphasis</em> FAIL — bold destroyed FAIL — leaked em PASS
4 _standalone emphasis_ <em>standalone emphasis</em> varies PASS PASS
5 **bold** and _emphasis_ <strong>bold</strong> and <em>emphasis</em> FAIL PASS PASS
6 get_value (2 parts, 1 underscore) get_value (literal) PASS PASS PASS
7 get_user_value (3 parts, 2 underscores — odd pairing) get_user_value (literal) PASS FAIL if + _em_ in same paragraph PASS
8 `code_with_underscores` (inline code) <code>code_with_underscores</code> PASS PASS PASS
9 __dunder__ method name <strong>dunder</strong> (bold) FAIL PASS PASS

Reproduction Script

import markdown2

test_cases = [
    (
        "Adjacent bold lines",
        "**Name**: Alice\n**Role**: Engineer\n**Team**: quality_assurance",
    ),
    (
        "Mid-word underscores only",
        "Call load_user_profile to fetch data.\nAlso try get_config_value.",
    ),
    (
        "Mid-word underscore + standalone emphasis (triggers Bug 2)",
        "Call load_user_profile to fetch data.\nUse _caution_ when calling it.",
    ),
    (
        "Mixed bold + emphasis + identifiers",
        "**Name**: quality_assurance.data_pipeline\n"
        "Call get_user_profile for **lookups**.\n"
        "See _note_ and **warning** formatting.",
    ),
]

for name, md_text in test_cases:
    md = markdown2.Markdown(extras={"middle-word-em": False})
    html = md.convert(md_text)
    has_broken_bold = "<em>*" in html
    has_leaked_em = any(
        c.isalnum() and html[i+1:i+5] == "<em>" and html[i+5:i+6].isalnum()
        for i, c in enumerate(html[:-6])
    )
    status = "FAIL" if (has_broken_bold or has_leaked_em) else "OK"
    print(f"{status}: {name}")
    if status == "FAIL":
        print(f"  HTML: {html.strip()[:200]}")

Expected: All FAIL on markdown2 2.5.4; All OK after applying the proposed run() fix.


Workaround

Until this is fixed upstream, consumers can monkey-patch MiddleWordEm.run() at runtime:

import re
import markdown2

_midword_underscore_re = re.compile(r'(?<=[a-zA-Z0-9])_(?=[a-zA-Z0-9])')
_hash_text_fn = markdown2._hash_text

def _fixed_mwe_run(self, text):
    if self.options['allowed']:
        return text
    if self.md.order < self.md.stage:
        def _hash_midword(match):
            substr = match.group(0)
            key = _hash_text_fn(substr)
            self.hash_table[key] = substr
            return key
        text = _midword_underscore_re.sub(_hash_midword, text)
    else:
        orig_text = ''
        while orig_text != text:
            orig_text = text
            for key, substr in self.hash_table.items():
                text = text.replace(key, substr)
    return text

markdown2.MiddleWordEm.run = _fixed_mwe_run

This must be called before any markdown2.Markdown() instances are created.


Environment

  • Python: 3.13
  • markdown2: 2.5.4
  • OS: Linux
Ngôn ngữ chính
Python
Star
2.8k
Fork
459
Merge trung bình
2 ngày 19 giờ
Pull request đã merge (30 ngày)
4

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của trentm/python-markdown2

Tất cả issue của trentm/python-markdown2

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.