Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

2.5.5 regression: strong emphasis adjacent to word/CJK chars fails when content starts or ends with punctuation

Đang mở
#688 4 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
58/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
python
Lĩnh vực
backend

Hướng nghiên cứu

Bắt đầu với Markdown._do_italics_and_bold() và GFMItalicAndBoldProcessor, so sánh đường dẫn biểu thức chính quy của 2.5.4 với đường dẫn bộ xử lý của 2.5.5. Chạy các bản tái hiện tối thiểu và dài hơn trên cả hai phiên bản, sau đó thêm phạm vi kiểm thử hồi quy cho thấy văn bản chữ-số hoặc CJK liền kề có dấu câu được kết xuất thành cùng một HTML nhấn mạnh mạnh như trong 2.5.4.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Bug

Summary

After upgrading from markdown2 2.5.4 to 2.5.5, strong emphasis sometimes stops parsing when all of these are true:

  • the **...** span is adjacent to alphanumeric or CJK text without surrounding spaces
  • the emphasized content starts or ends with punctuation, for example Chinese quotes or parentheses

On 2.5.4 these cases render correctly. On 2.5.5 the parser leaves literal ** in the output, and in longer inputs it can also emit malformed mixed HTML.

This reproduces with the default parser configuration, no extras required.

Related to #679 because it also looks like an emphasis-regression in the recent parser changes, but this one reproduces without middle-word-em and with plain markdown2.markdown(...).

Minimal repro

import markdown2

cases = [
    'a**“b”**c',
    '**“b”**c',
    'a**(b)**c',
]

print('version:', markdown2.__version__)
for text in cases:
    print('INPUT :', text)
    print('OUTPUT:', markdown2.markdown(text).strip())
    print()
2.5.4
<p>a<strong>“b”</strong>c</p>
<p><strong>“b”</strong>c</p>
<p>a<strong>(b)</strong>c</p>
2.5.5
<p>a**“b”**c</p>
<p>**“b”**c</p>
<p>a**(b)**c</p>

Longer repro that produces malformed mixed HTML

import markdown2

text = '*   **示例**:系统会使用**(方案A)**或**(方案B)**进行处理。'
print(markdown2.markdown(text))
2.5.4
<ul>
<li><strong>示例</strong>:系统会使用<strong>(方案A)</strong>或<strong>(方案B)</strong>进行处理。</li>
</ul>
2.5.5
<ul>
<li><strong>示例</strong>:系统会使用**(方案A)<strong>或</strong>(方案B)**进行处理。</li>
</ul>

Suspected regression window

Looking at the source, Markdown._do_italics_and_bold() changed between these versions:

2.5.4
text = self._strong_re.sub(r"<strong>\\2</strong>", text)
text = self._em_re.sub(r"<em>\\2</em>", text)
2.5.5
if not self._iab_processor:
    self._iab_processor = GFMItalicAndBoldProcessor(self, None)
if self._iab_processor.test(text):
    text = self._iab_processor.run(text)

So this looks related to the switch to GFMItalicAndBoldProcessor as the default implementation for _do_italics_and_bold().

Environment

  • Python: 3.14.0
  • markdown2: 2.5.4 vs 2.5.5
  • OS: Windows 11
Ngôn ngữ chính
Python
Star
2.8k
Fork
459
Merge trung bình
2 ngày 19 giờ
Pull request đã merge (30 ngày)
4

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của trentm/python-markdown2

Tất cả issue của trentm/python-markdown2

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.