RTF reader does not combine UTF-16 surrogate pairs

#11920 · open · 0 comments

View on GitHub ↗

VijayVignesh1

<!-- Thank you for reporting an issue! Before you continue, please make sure that you have - reproduced your issue with the latest release of pandoc (https://github.com/jgm/pandoc/releases), or online (https://pandoc.org/try). - searched the issue tracker (https://github.com/jgm/pandoc/issues) for similar issues (including closed issues). - searched the discussion forum (https://github.com/jgm/pandoc/discussions) for solutions. If you are experiencing a regression (an unwanted change of behavior from an earlier version of pandoc), please indicate this and give relevant version numbers. Be sure to check the release notes (https://pandoc.org/releases) for relevant changes, as you might find a solution there. Note that this bug tracker is for reporting bugs, not asking questions. For questions, use the discussion forum: https://github.com/jgm/pandoc/discussions. --> **Explain the problem.** The RTF reader does not correctly decode a UTF-16 surrogate pair represented using consecutive \uN control words. For example: `\uc1\u-10187 ?\u-9157 ?` Expected output is a latin character 𝐻 Actual output is the fallback character: �� Sample Input: [sample.rtf](https://github.com/user-attachments/files/32921438/sample.rtf) Command used: ``` pandoc -f rtf -t html sample.rtf -o sample.html ``` Sample Output: [sample.html](https://github.com/user-attachments/files/32921439/sample.html) **Pandoc version?** Pandoc version: 3.12 OS: Ubuntu 22.04.5 LTS (Jammy Jellyfish), x86_64

Comments