char-width

npm

2 min read Original article โ†—

1.0.0 โ€ข Public โ€ข Published

char-width

npm

A TypeScript/JavaScript port of Rust's unicode-width: sequence-aware Unicode display width for terminals, with O(1) work per code point.

Attention to Accuracy

  • ๐ŸŽฏ Verified against every output of Rust's unicode-width crate
  • ๐Ÿงฌ Sequence-aware: flags, keycaps, emoji ZWJ and presentation/modifier sequences, etc.
  • ๐ŸŒ Unicode 17.0, with a cjk mode for East Asian (ambiguous-wide) contexts

Getting Started

Install char-width via npm:

Usage

import { charWidth, strWidth } from 'char-width';

charWidth('a'); // 1
charWidth('ๅฅฝ'); // 2
charWidth('\x1b'); // undefined (control character)

strWidth('hello'); // 5
strWidth('โค๏ธ'); // 2 (emoji presentation sequence)
strWidth('๐Ÿ‘ฉโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆ'); // 2 (ZWJ sequence)
strWidth('๐Ÿ‡ฆ๐Ÿ‡บ'); // 2 (flag)

Function Parameters:

charWidth():

  • char: The string whose first code point is measured.
  • cjk: Optional; treats East Asian Ambiguous characters as wide.

strWidth():

  • str: The string to measure.
  • cjk: Optional; treats East Asian Ambiguous characters as wide.

Scope of Output

charWidth():

  • 0, 1, 2, or 3: width of the first code point of char.
  • undefined: when it's a control character (C0, DEL, C1) or the string is empty.

strWidth():

  • non-negative width of str: control characters and "\r\n" each count as width 1.

Documentation

The behavior exactly follows that of the unicode-width crate โ€” see its documentation.

TL;DR:

  • A character's width can depend on what follows it (VS16, ZWJ), so strWidth scans once, back to front โ€” O(1) work per code point.
  • Canonically equivalent strings get the same width.
  • Widths predict terminal behavior; no library matches every terminal.

โš ๏ธ One deliberate divergence

A lone surrogate โ€” possible in a JS string, impossible in a Rust one โ€” has width 1.

Updating to a new Unicode version

git clone https://github.com/unicode-rs/unicode-width # reference crate
(cd unicode-width && python3 scripts/unicode.py)      # regenerate Rust tables from the UCD
npm run gen                                           # convert them to src/gen/*.ts
npm run gen:truth                                     # re-dump ground truth (needs cargo)
npm test                                              # 1.1M-code-point conformance check

Feedback

Found something odd?
Feel free to open an issue.

Acknowledgments

  • Width algorithm, state machine, and tables derived from unicode-width by the Rust Project Developers and the unicode-rs maintainers (MIT OR Apache-2.0, used under the MIT option).
  • Character data from the Unicode Character Database (Unicode License v3).

See THIRD-PARTY-NOTICES.md for more.

License

Distributed under the MIT License. See LICENSE for more information.

Looking for a POSIX-compliant port?

Try wcwidth-o1.

Readme

Keywords