Files
Jyotisha/scripts/fonts/CHARSET_SOURCE.txt
T
Jesse_ChenandClaude Opus 5.5 308f5502cb feat(fonts): self-hosted Jyotisha Serif SC heading slices (T1)
Renamed, sliced subset of Noto Serif SC SemiBold (OFL 1.1, noto-cjk
Serif2.003) covering the 6500 level-1 + level-2 characters of the
通用规范汉字表 plus Latin, punctuation and common symbols.

- scripts/fonts/build_serif_slices.py: fontTools subsetting into 30 woff2
  slices + serif-sc.css (@font-face, weight 500 700, font-display swap);
  byte-identical output for the same input.
- Frequency order (char_order.txt): repo corpus top 300, then Google
  Fonts SC frequency bands (nam-files, Apache-2.0), then corpus count;
  unranked characters sliced by contiguous code point.
- tests/test_serif_font_slices.py: coverage, no overlap, files exist,
  120 KB cap; added to the quick quality gate.

TASK-serif-headings-20260928 T1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199rbQDTsUbCVw84wc8BTFe
2026-09-28 19:31:43 +08:00

38 lines
2.1 KiB
Plaintext

《通用规范汉字表》(2013) character lists used by build_serif_slices.py
tygfhzb-level-1.txt 一级字表 3500 characters, one per line
tygfhzb-level-2.txt 二级字表 3000 characters, one per line
Official publication: 国务院《通用规范汉字表》(国发〔2013〕23号), PDF
http://www.gov.cn/gzdt/att/att/site1/20130819/tygfhzb.pdf
Machine-readable copy taken verbatim from
https://github.com/shengdoushi/common-standard-chinese-characters-table
commit d9b599a9c9cc0dd2d58cad829e285bc780cd4451, files level-1.txt / level-2.txt
sha256 level-1.txt 79a6c710013cc86617d5db65871f59b2d67dee72415d380a2cc7145a51450fe4
sha256 level-2.txt d597a6e99ea7b8c41215081f824c6e8587bf1c9a3f45fb13086a826306789785
(the files here are byte-identical to those; the build script asserts
6500 distinct single characters)
Order: the table lists characters by stroke count, not by frequency.
char_order.txt holds the frequency order used for slicing (columns: character,
Google Fonts band, corpus count). Sort key:
1. the 300 characters most used in this repo's corpus (git-tracked
frontend/src, frontend/docs, references/, SKILL.md), so the astrology
vocabulary of headings sits in the first slices;
2. then the Google Fonts Simplified Chinese frequency band (0 = most
frequent; 99 = not in the 20 frequency bands);
3. then corpus count, then table order.
Google Fonts frequency bands (derived data only; the source file is not committed):
https://github.com/googlefonts/nam-files slices/simplified-chinese_default.txt
commit 2a68014b16056e965c18fd47de1f3fde9a5b0095
URL https://raw.githubusercontent.com/googlefonts/nam-files/2a68014b16056e965c18fd47de1f3fde9a5b0095/slices/simplified-chinese_default.txt
sha256 cd022036048744a478b2d261131a90400bacbb9098b83efd0360f1e44122c2c1
License: Apache License 2.0 (Copyright Google LLC). This is the published
CJK slicing data of the Google Fonts project; its first 20 subsets are
"FreqRange" bands in descending frequency (the file header says so).
Regenerate:
python3 scripts/fonts/build_serif_slices.py --refresh-order --gf-slices <simplified-chinese_default.txt>