Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Encode canonical charset names differ from Perl, breaking IO::HTML

Open
#1,445 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
65/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
java, perl
Domain
backend

Research direction

Start with the PerlOnJava Encode implementation and its existing compatibility tests, then reproduce the canonical-name failures on both the JVM and interpreter backends. Add focused regression coverage for the ISO-8859-15 and cp1252 mappings, and verify the IO::HTML 1.004 suite passes on both backends.

Written by the indexing model from the issue text.

Description

area:cpan-port area:unicode bug

Summary

IO::HTML 1.004 passes its complete upstream test suite under system Perl but fails 23 of 165 tests under PerlOnJava because Encode::find_encoding(...)->name returns Java-style charset names instead of Perl's canonical names.

Reproduction

CPAN random tester run: 20260918-141920-96054

Distribution: IO::HTML 1.004

System Perl 5.42.2:

All tests successful.
Files=5, Tests=165
Result: PASS

PerlOnJava JVM backend:

Files=5, Tests=165
Result: FAIL
Failed 2/5 test programs. 23/165 subtests failed.

The interpreter backend reproduces the same 23 failures.

Failing tests

  • t/10-find.t: 14 failures
  • t/20-open.t: 9 failures
  • t/00-all_prereqs.t, t/00-load.t, and t/30-outfile.t: pass

The failures are all charset-name mismatches:

got:      ISO-8859-15
expected: iso-8859-15

got:      windows-1252
expected: cp1252

The affected inputs include HTML declarations for ISO-8859-15 and ISO-8859-1, including declarations near the 1024-byte scan boundary.

Likely cause

IO::HTML calls Encode::find_encoding($charset) and then uses the returned object's name method. Perl returns Encode's canonical names (iso-8859-15 and cp1252), while PerlOnJava returns Java canonical charset names (ISO-8859-15 and windows-1252).

The PerlOnJava Encode implementation already maps US-ASCII to ascii and UTF-8 to utf-8-strict, but otherwise returns the Java charset name unchanged. The canonical-name mapping needs to cover the relevant Perl aliases and should preserve the expected behavior for both backends.

Expected behavior

For equivalent Encode::find_encoding calls, PerlOnJava should return an Encode::Encoding object whose name matches Perl's canonical name, including at least:

  • ISO-8859-15 -> iso-8859-15
  • windows-1252 / ISO-8859-1 -> cp1252

Please add focused regression coverage for the canonical names and verify IO::HTML 1.004 on both the JVM and interpreter backends after the fix.

Related project context

The project module-compatibility notes already identify encoding-name case/alias normalization as an outstanding Encode compatibility gap affecting IO::HTML. This report provides a complete upstream reproduction and confirms the defect on both PerlOnJava execution backends.

Dominant language
Perl
Stars
64
Forks
6
Avg merge
5h 11m
Merged PRs (30d)
168

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from fglock/PerlOnJava

All issues in fglock/PerlOnJava

Similar issues

More Perl issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.