Renamed, sliced subset of Noto Serif SC SemiBold (OFL 1.1, noto-cjk Serif2.003) covering the 6500 level-1 + level-2 characters of the 通用规范汉字表 plus Latin, punctuation and common symbols. - scripts/fonts/build_serif_slices.py: fontTools subsetting into 30 woff2 slices + serif-sc.css (@font-face, weight 500 700, font-display swap); byte-identical output for the same input. - Frequency order (char_order.txt): repo corpus top 300, then Google Fonts SC frequency bands (nam-files, Apache-2.0), then corpus count; unranked characters sliced by contiguous code point. - tests/test_serif_font_slices.py: coverage, no overlap, files exist, 120 KB cap; added to the quick quality gate. TASK-serif-headings-20260928 T1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0199rbQDTsUbCVw84wc8BTFe
38 lines
2.1 KiB
Plaintext
38 lines
2.1 KiB
Plaintext
《通用规范汉字表》(2013) character lists used by build_serif_slices.py
|
|
|
|
tygfhzb-level-1.txt 一级字表 3500 characters, one per line
|
|
tygfhzb-level-2.txt 二级字表 3000 characters, one per line
|
|
|
|
Official publication: 国务院《通用规范汉字表》(国发〔2013〕23号), PDF
|
|
http://www.gov.cn/gzdt/att/att/site1/20130819/tygfhzb.pdf
|
|
|
|
Machine-readable copy taken verbatim from
|
|
https://github.com/shengdoushi/common-standard-chinese-characters-table
|
|
commit d9b599a9c9cc0dd2d58cad829e285bc780cd4451, files level-1.txt / level-2.txt
|
|
sha256 level-1.txt 79a6c710013cc86617d5db65871f59b2d67dee72415d380a2cc7145a51450fe4
|
|
sha256 level-2.txt d597a6e99ea7b8c41215081f824c6e8587bf1c9a3f45fb13086a826306789785
|
|
(the files here are byte-identical to those; the build script asserts
|
|
6500 distinct single characters)
|
|
|
|
Order: the table lists characters by stroke count, not by frequency.
|
|
char_order.txt holds the frequency order used for slicing (columns: character,
|
|
Google Fonts band, corpus count). Sort key:
|
|
1. the 300 characters most used in this repo's corpus (git-tracked
|
|
frontend/src, frontend/docs, references/, SKILL.md), so the astrology
|
|
vocabulary of headings sits in the first slices;
|
|
2. then the Google Fonts Simplified Chinese frequency band (0 = most
|
|
frequent; 99 = not in the 20 frequency bands);
|
|
3. then corpus count, then table order.
|
|
|
|
Google Fonts frequency bands (derived data only; the source file is not committed):
|
|
https://github.com/googlefonts/nam-files slices/simplified-chinese_default.txt
|
|
commit 2a68014b16056e965c18fd47de1f3fde9a5b0095
|
|
URL https://raw.githubusercontent.com/googlefonts/nam-files/2a68014b16056e965c18fd47de1f3fde9a5b0095/slices/simplified-chinese_default.txt
|
|
sha256 cd022036048744a478b2d261131a90400bacbb9098b83efd0360f1e44122c2c1
|
|
License: Apache License 2.0 (Copyright Google LLC). This is the published
|
|
CJK slicing data of the Google Fonts project; its first 20 subsets are
|
|
"FreqRange" bands in descending frequency (the file header says so).
|
|
|
|
Regenerate:
|
|
python3 scripts/fonts/build_serif_slices.py --refresh-order --gf-slices <simplified-chinese_default.txt>
|