rubrapack Manual←↑→

23 Text: characters, code points and encodings

A file holds bytes, and a byte is a number (chapter 1). Text - a product name, a file name, a license - is stored by giving every character a number and writing the numbers as bytes. There are several ways to do both, and a package that uses the wrong one shows garbage. This chapter explains the ways rubrapack and Windows use.

23.1 Characters and code points#

Unicode is the world's catalogue of characters: letters of every script, digits, punctuation, symbols, emoji - about 160,000 of them. Each has a number, its code point, written U+ and at least four hex digits:

CharacterCode pointName in the catalogue
HU+0048LATIN CAPITAL LETTER H
Γ©U+00E9LATIN SMALL LETTER E WITH ACUTE
€U+20ACEURO SIGN
ν•œU+D55CHANGUL SYLLABLE HAN
πŸ˜€U+1F600GRINNING FACE

Code points go up to U+10FFFF. A code point says which character; it does not yet say which bytes to write. That is the job of an encoding.

23.2 ASCII#

The oldest encoding still everywhere is ASCII (1963): 128 characters, code points 0 to 127, one byte each. It has the English letters, digits, punctuation and some control characters:

BytesCharacters
00-1fcontrol characters: 09 tab, 0a line feed (new line), 0d carriage return
20space
30-390-9
41-5aA-Z
61-7aa-z
2e 2f 3a 5c. / : \

The tutorial's dist\docs\guide.txt holds six bytes:

47 75 69 64 65 0a
 G  u  i  d  e  (new line)

The first 128 code points of Unicode are ASCII's, and every encoding below writes them as the same single byte. That is why English text looks the same in almost any encoding - and why problems show only with other letters.

Windows ends a line with two bytes, 0d 0a (carriage return, line feed); Linux and macOS with 0a alone. Most programs read both.

23.3 Code pages: one byte, 256 characters#

ASCII leaves the values 128 to 255 of a byte unused. For decades each language region filled them differently, and each filling is a code page, numbered by Microsoft:

Code pageRegionΓ©β‚¬ν•œ
1252Western Europee980-
949Korean-a2 e6c7 d1
65001all: this is UTF-8c3 a9e2 82 aced 95 9c

A code page can hold only its own region's characters (949 uses two bytes for Korean). And bytes mean different things in different code pages: the byte e9 is Γ© in 1252 but half of a Korean character in 949. Text written in one and read in another comes out as garbage, known by the Japanese word mojibake.

Every MSI package has one code page for all its text. rubrapack always writes 65001, UTF-8, so any language fits into one package; tutorial chapter 7 shows inspect reporting it. Part IV shows a trap: the summary information has a code page of its own.

23.4 UTF-8#

UTF-8 writes every Unicode code point in one to four bytes. ASCII stays one byte, so an ASCII file is already UTF-8. Larger code points use more bytes, following a fixed bit pattern:

Code pointsBytesBit pattern (x = bits of the code point)
U+0000 - U+007F10xxxxxxx
U+0080 - U+07FF2110xxxxx 10xxxxxx
U+0800 - U+FFFF31110xxxx 10xxxxxx 10xxxxxx
U+10000 - U+10FFFF411110xxx 10xxxxxx 10xxxxxx 10xxxxxx

Worked out for ν•œ, U+D55C:

  D55C in binary:       1101 0101 0101 1100      (16 bits)
  split 4 + 6 + 6:      1101   010101   011100
  into the pattern:  1110 1101  10 010101  10 011100
  in hex:                ED         95         9C

So ν•œ is ed 95 9c in UTF-8. The pattern makes UTF-8 easy to check: a byte starting 10 is never the first byte of a character, so a reader can always find where characters begin, and bytes of another encoding almost never form valid UTF-8 by accident. rubrapack stops at the first byte of a source that breaks the pattern (RP1001: not valid UTF-8) rather than guessing.

23.5 UTF-16#

UTF-16 writes code points up to U+FFFF as one 2-byte unit, and the rest as two units (a surrogate pair). Windows uses it inside: file names, the registry, and every text in a Compound File's directory are UTF-16, little-endian (chapter 2):

  H     e     l     l     o
  48 00 65 00 6c 00 6c 00 6f 00        "Hello" in UTF-16LE: ASCII letters get a 00 after them

  ν•œ
  5c d5                                U+D55C: the two bytes of D55C, little end first

  πŸ˜€
  3d d8 00 de                          U+1F600: too big for one unit, so the pair D83D DE00

A text in a dump that looks like letters with dots between them - H.e.l.l.o. - is UTF-16.

23.6 The byte order mark#

A file may start with the code point U+FEFF, the byte order mark (BOM), to say which encoding it is: ef bb bf for UTF-8, ff fe for UTF-16LE. Windows Notepad used to add one to UTF-8 files. rubrapack build and lint read a source in UTF-8, with or without a BOM, and in UTF-16 with its BOM; edit changes UTF-8 sources only.

23.7 Unicode normalization#

Some characters can be written in Unicode in more than one way. Γ© is one code point, U+00E9 - or two: the letter e, U+0065, followed by U+0301, a combining acute accent that sits on the letter before it. Both look exactly the same on screen:

FormCode pointsUTF-8 bytes
composed (NFC)U+00E9c3 a9
decomposed (NFD)U+0065 U+030165 cc 81

Korean is the largest case. Every modern Hangul syllable has a composed code point, and can also be written as its letters (jamo):

FormCode pointsUTF-8 bytes
NFCU+D55C ν•œed 95 9c
NFDU+1112 α„’ U+1161 α…‘ U+11AB ᆫe1 84 92 e1 85 a1 e1 86 ab

(The composed code point is computed from the letters: 0xAC00 + (18 x 21 + 0) x 28 + 4 = 0xD55C, where 18, 0 and 4 are the numbers of γ…Ž, ㅏ and γ„΄ in the Unicode tables.)

Normalization turns text into one agreed form. NFC ("composed") is the form Windows, the web and almost every keyboard produce. NFD ("decomposed") is what macOS file systems store file names in - so a file copied from a Mac may have a name whose bytes differ from the same name typed on Windows. For Windows the two are different names: a package could then install ν•œ.txt twice, or a shortcut could miss its target.

That is why rubrapack has build --nfc (tutorial chapter 15), which turns file, folder and shortcut names into NFC, and why lint warns (RP2105) about text that is not in NFC (tutorial chapter 18). rubrapack's normalization follows the Unicode 17.0 tables.

23.8 When text looks wrong#

You seeLikely cause
é instead of Γ©UTF-8 bytes c3 a9 read as code page 1252
οΏ½λΈ³ or ? instead of ν•œUTF-8 bytes read as code page 949, or the other way round
H.e.l.l.o. in a dumpUTF-16 read as bytes
two names that look the same but are notone NFC, one NFD

23.9 Where this is used#