
CHAPTER 14 / COMPRESSION OF UNICODE FILES
293
Table 14.3 UCS-2 and UTF-8 Coding
Data Bits Input Bit Pattern (UCS-2)
Coding into Successive Bytes (UTF-8)
7
11
16
0000 0000 Oabc defg
0000 Oabc defg hijk
abcd efgh ijkl mnop
Oabc defg
llOa bcde lOfg hijk
iii0 abcd lOef ghij lOkl mop
A text string in UCS-2 may, when converted to UTF-8,
1. contract if it contains mostly ASCII characters,
2. remain essentially the same size if it is mostly Cyrillic, Hebrew, or Arabic, or
3. expand if it is predominately some Asian alphabet.
14.3 COMPRESSION OF UNICODE
A Unicode file is a sequence of characters in either UCS-2 or UTF-8 ...