Invisible character code points and Unicode names
The notation U+ followed by hexadecimal digits identifies a Unicode code point. Names below describe characters, not their author or source. The appropriate cleanup action depends on where the character occurs and what the destination expects.
- U+200B ZERO WIDTH SPACE: an invisible line-break opportunity; inspect unexpected occurrences inside exact-match values.
- U+200C ZERO WIDTH NON-JOINER and U+200D ZERO WIDTH JOINER: joining controls; preserve meaningful language and emoji sequences.
- U+2060 WORD JOINER: suppresses a line break without adding visible space.
- U+FEFF: used as a byte order mark in an encoding context; review an unexpected occurrence inside text.
- U+00A0 NO-BREAK SPACE and U+202F NARROW NO-BREAK SPACE: space characters with no-break behavior; replace only when ordinary spacing is intended.
- U+00AD SOFT HYPHEN: a discretionary hyphenation character; distinguish it from a visible hyphen.
- U+200E, U+200F, and directional controls in U+202A–U+202E and U+2066–U+2069: affect bidirectional layout; review the surrounding text.
- U+FE0E and U+FE0F: presentation selectors used with eligible characters; removing them can change appearance.
How to interpret character counts and positions
A count tells you how many covered code points the scanner observed. It does not tell you how many visible symbols are wrong, how severe the problem is, or whether all occurrences should receive the same treatment. Several characters can combine into one displayed symbol, while one formatting character can influence a larger span.
Look at location as well as count. An unexpected character inside a database key deserves a different response from the same character in a demonstration of Unicode behavior. The intended content and the destination’s rules determine whether to preserve, remove, or replace it.
Remove or replace: choosing the correct action
Deleting a space-like character between two words can join those words. Replacing it with U+0020 preserves separation while changing line-breaking behavior. Removing a zero-width artifact inside an identifier is a different operation with a different reason.
Punctuation conversion is another distinct choice. A curly quotation mark or an em dash is visible, meaningful text, even when a particular destination expects simpler typography. Keep those transformations explicit so you can explain why the output differs from the original.
Unicode characters beyond this reference list
Unicode includes combining marks, controls, variation sequences, and script-specific behavior beyond this list. A tool can only report the patterns it implements. “Nothing found” should be read as a statement about that coverage, not as proof that the entire input has no unusual properties.
If your issue persists, inspect the text in a code-point-aware editor and consult the relevant standard or destination documentation. Normalization differences and visually similar letters can cause mismatches without involving the common invisible characters listed here. Removing more characters at random is unlikely to be a reliable diagnosis.
Document and verify character corrections
For a repeatable fix, save the original sample, name the code point, state the intended replacement, and record the result of the destination check. You can represent the character with a visible label such as “[U+200B]” in a bug report rather than relying on a character readers cannot see.
When language or emoji is involved, compare the rendered output as well as the underlying sequence. A successful cleanup preserves the content you meant to keep while changing the specific artifact you meant to fix. That is a stronger standard than simply producing the smallest possible character count.
Sources & further reading
Primary references for the technical points in this guide.