Go library for detecting and fixing Cyrillic text encoding issues (mojibake).
Handles common encoding problems with Russian and other Cyrillic text:
| Problem | Function | Example |
|---|---|---|
| CP1251 bytes stored as ISO-8859-1 | ReencodeISO8859AsCP1251 |
Latin gibberish -> correct Cyrillic |
| UTF-8 misread as ISO-8859-1 | DecodeMojibakeFromISO8859 |
Garbled accented chars -> correct UTF-8 |
| UTF-8 misread as CP1251 (double-encoded) | DecodeMojibakeFromCP1251 |
Multi-byte garble -> correct UTF-8 |
go get github.com/drgolem/cyrillic-encodingpackage main
import (
"fmt"
cyrillic "github.com/drgolem/cyrillic-encoding"
)
func main() {
// Fix CP1251 bytes that were decoded as ISO-8859-1
fixed, ok := cyrillic.ReencodeISO8859AsCP1251(garbledText)
if ok {
fmt.Println(fixed) // correct Cyrillic text
}
// Fix UTF-8 text that was misread as CP1251 (double-encoded mojibake)
fixed = cyrillic.DecodeMojibakeFromCP1251(mojibakeText)
// Fix UTF-8 text that was misread as ISO-8859-1
fixed = cyrillic.DecodeMojibakeFromISO8859(mojibakeText)
// Check if text is pure printable ASCII (no high bytes)
if cyrillic.IsASCIIPrintable(text) {
fmt.Println("text is ASCII-only")
}
// Count Cyrillic characters (weighted: lowercase=2, uppercase=1)
score := cyrillic.CountCyrillic(text)
}ReencodeISO8859AsCP1251(str string) (string, bool)- Fix CP1251 bytes decoded as ISO-8859-1. Returns the corrected string andtrueon success.DecodeMojibakeFromISO8859(mojibake string) string- Fix UTF-8 bytes misread as ISO-8859-1.DecodeMojibakeFromCP1251(mojibake string) string- Fix UTF-8 bytes misread as CP1251. Uses Cyrillic character heuristics to validate.
IsASCIIPrintable(str string) bool- Check if all characters are printable ASCII (bytes 32-127).CountCyrillic(s string) int- Count Cyrillic characters with weighted scoring (lowercase=2x).CP1251ToByte(r rune) byte- Convert a Unicode rune to its CP1251 byte value.
This package has zero external dependencies - only Go standard library.
MIT