GPTClean-up

Debugging example

Zero-Width Space in Code: How to Find and Remove U+200B

A zero-width space in code can create an unexpected character inside an identifier, string, or configuration key. Learn how to identify U+200B, compare a failing value with a known-good sample, and verify a targeted removal without changing intentional Unicode.

Example: U+200B inside a code identifier or key

Imagine an application expects the key “accountId”, but a copied key contains “account[U+200B]Id”. The bracketed code-point label is shown here for explanation; the real character may have no visible width. The two sequences are different even when they look the same in a normal editor.

Do not assume the language will handle the extra character the same way in every position. A key inside a string, an identifier in source, and whitespace between tokens are distinct contexts. Test the actual failing input instead of relying on a generic claim that every invisible character always causes a syntax error.

Compare typed and copied code values

Create the shortest example that still fails. Type the expected value manually and compare it with the pasted value in the same environment. If one works and the other does not, inspect the character sequence and record the exact difference.

Keep this comparison small. Replacing a whole file makes it harder to know which change mattered and can alter line endings or unrelated content. A one-character correction provides much clearer evidence when you rerun the original failing operation.

Remove an unwanted zero-width space

An unexpected U+200B inside an exact-match key may be a reasonable removal target. The same code point in a language sample or a test fixture may be intentional. Check the surrounding code and the expected behavior before applying the edit.

If your editor can show hidden characters, use that view during review. When documenting the issue, write the escaped form or the U+200B label rather than depending on readers to see the invisible character. This makes bug reports and code-review comments easier to understand.

Test code after removing U+200B

After the patch, review the diff and run the check that originally failed. If the issue was a parsed key, inspect the parsed value. If it was source syntax, use the parser and the relevant focused tests. Do not stop at a cleaner’s success message.

If the problem remains, consider a different character, a normalization difference, a case mismatch, or a separate code issue. The absence of zero-width spaces does not certify that the rest of the program is valid or safe.

Prevent zero-width spaces in copied code

Trace recurring cases through the applications used to copy, format, and publish the snippet. Compare the source before rendering with the copied result where you can. An editor warning or corrected documentation export may prevent the issue more reliably than repeated cleanup.

Preserve intentional Unicode and avoid a repository-wide “remove all non-ASCII” rule. A good prevention rule is narrow enough to catch the known failure while leaving valid language text and test data alone. The aim is predictable code handling, not an artificially restricted character set.

Sources & further reading

Primary references for the technical points in this guide.

Frequently Asked Questions

Can a zero-width space be valid text?

Yes. It has line-breaking uses, and it may also be intentional test data. The code context determines whether it is unwanted.

Why use escaped notation in a bug report?

It makes the exact character visible and avoids accidentally copying an unexplained invisible character into another file.

Keep exploring

Related Tools and Guides