WIPIVERSE

GB 2312

GB 2312 (officially GB/T 2312-1980) is a key official character set of the People's Republic of China, used for encoding Simplified Chinese characters. The designation "GB" refers to Guobiao (国家标准), meaning "national standard," while the "T" suffix (推荐; tuījiàn) denotes a recommended (non-mandatory) standard. GB2312 is also the registered internet name for EUC-CN, which is its usual encoded form.

History and Status

GB/T 2312-1980 was originally a mandatory national standard designated GB 2312-1980. However, following a National Standard Bulletin of the People's Republic of China in 2017, GB 2312 is no longer mandatory, and its standard code was modified to GB/T 2312-1980. The standard has been superseded by GBK and GB 18030, which include additional characters, but GB/T 2312 remains in widespread use as a subset of those encodings.

Character Coverage

The GB/T 2312 standard includes 6,763 Chinese characters (arranged in two levels) along with symbols, punctuation, Japanese kana, the Greek and Cyrillic alphabets, Zhuyin (Bopomofo), and a double-byte set of Pinyin letters with tone marks. In the later version GB/T 2312-1980, there are 7,445 characters total, comprising 682 signs and 6,763 Chinese characters.

Characters are arranged in a 94×94 grid (as in ISO 2022), with each two-byte code point expressed in the qūwèi (区位) form, which specifies a row (区; qū) and a position within the row (位; wèi). The rows are organized as follows:

  • Rows 01–09: punctuation, special characters, Hiragana, Katakana, Greek, Cyrillic, Pinyin, and Bopomofo
  • Rows 16–55: first level of Chinese characters, arranged by Pinyin (3,755 characters)
  • Rows 56–87: second level of Chinese characters, arranged by radical and stroke count (3,008 characters)
  • Rows 10–15 and 88–94: unassigned

Encoding Forms

GB/T 2312 can be encoded in several formats:

EUC-CN: The most common encoding form, using two bytes per character not found in ASCII. The first byte ranges from 0xA1–0xF7 and the second from 0xA1–0xFE. Compared to UTF-8, which uses three bytes per CJK ideograph, GB/T 2312 uses only two, making it more storage-efficient, though it covers fewer ideographs than Unicode.

ISO-2022-CN: The encoding specified in the official documentation, which uses the same byte range as ASCII (0x21–0x77 for the first byte, 0x21–0x7E for the second) and requires escape sequences to switch between ASCII and the two-byte character range.

HZ: An encoding used mostly for Usenet postings, using the same byte pairs as ISO-2022-CN but with different escape sequences.

Usage and Legacy

As of July 2026, GB2312 is the second-most popular encoding served from China and its territories (after UTF-8), with 3.3% of web servers serving a page declaring it. Globally, GB2312 is declared on less than 0.04% of all web pages. All major web browsers decode GB2312-marked documents as if they were marked with the superset GBK encoding, except for Safari and Edge on the label "GB_2312."

There is an analogous character set known as GB/T 12345, which supplements GB/T 2312 with traditional character forms by replacing simplified forms in their qūwèi code, along with some additional supplemental characters. GB-encoded fonts often come in pairs, one with the GB/T 2312 (simplified) character set and the other with the GB/T 12345 (traditional) character set.

Related Standards

GB/T 2312 is related to other CJK national character set standards, including JIS X 0208 (Japan) and KS X 1001 (South Korea), which share the same ISO 2022-based 94×94 grid structure. It has been succeeded by GBK and GB 18030, the latter being the current mandatory Unicode-compatible Chinese national standard.

Browse

More topics to explore

    Browse all articles