Standardizing Full-Width Alphanumeric Characters with PowerShell Unicode Normalization NFKC – A Non-Destructive Method for Source Data

PowerShellカテゴリを表すパンダのイラスト PowerShell

About this article
This article was created using an automated generation workflow leveraging generative AI. We verified the official specifications of .NET String.Normalize and NormalizationForm.FormKC and created a sample to observe Unicode normalization via PowerShell. Execution verification on a physical PowerShell environment has not been performed.

Verification Status: 📘 .NET Official Specifications Verified / PowerShell Physical Environment Unverified

When 'ABC123' (full-width) and 'ABC123' (half-width) are mixed in Excel or CSV files, they may appear nearly identical but be compared as distinct values. The same applies when half-width katakana or encircled numbers are mixed in.

In .NET,String.NormalizewithNormalizationForm.FormKCspecified,Unicode compatibility normalization (NFKC)can unify many variations in notation. However, because NFKC includes conversions that cannot restore the original representation, it must not be applied unconditionally to product codes, employee IDs, passwords, and similar data.

Trying a Minimal Example in PowerShell

We assume an environment where .NETString.Normalizeis available, such as PowerShell 7 or Windows PowerShell 5.1.

$before = 'ABC123'
$after  = $before.Normalize([Text.NormalizationForm]::FormKC)

"Before: $before"
"After : $after"

The expected display is as follows. This is a result anticipated from official specifications and is not physical machine output.

Before: ABC123
After : ABC123

FormKCis NFKC normalization, which performs compatibility decomposition and canonical composition.FormC(NFC), which is very similar, primarily handles canonical equivalence and has a different purpose from unifying compatibility characters.

Comparing Various Characters

The following sample presents a table before and after conversion to verify whether the content actually changed. Input consists solely of dummy strings and does not write to files or the registry.

$values = @(
    'ABC123'
    'カタカナ'
    '① Ⅳ'
)

$results = foreach ($original in $values) {
    $normalized = $original.Normalize([Text.NormalizationForm]::FormKC)

    [pscustomobject]@{
        Before  = $original
        After   = $normalized
        Changed = ($original -cne $normalized)
    }
}

$results | Format-Table -AutoSize

The expected conversion isABC123→ABC123、カタカナ→カタカナ、① Ⅳ→1 IV. You can also see this from the changes in the representation of circled numbers and Roman numerals, which indicates that this isnot simply deleting spaces or changing the width of alphanumeric characters.

The same code isstored in the Daily-Code-Samples NFKC experiment. The sample only reads and displays data, and does not automatically overwrite CSVs.

Change one location to compare with NFC

FormKCtoFormCand compare the same characters.

$text = 'ABC123'
$text.Normalize([Text.NormalizationForm]::FormC)
$text.Normalize([Text.NormalizationForm]::FormKC)

FormCWhile full-width alphanumeric characters are expected to remain unchanged inFormKC, they are expected to be standardized to half-width alphanumeric characters in . Rather than one being the "correct" choice, it depends onwhat you want to treat as the same character

. The official .NETString.NormalizeandNormalizationFormdocumentation provides explanations of normalization forms.

When using this for work, keep the "pre-normalized" data

For example, if product name fields in sales data contain a mix of full-width and half-width characters depending on the person entering them, it is useful to add a separate normalized column for searching and aggregation.

$rows = @(
    [pscustomobject]@{ Item = 'ABC123'; Count = 2 }
    [pscustomobject]@{ Item = 'ABC123';   Count = 3 }
)

$rows | Select-Object Item, Count, @{
    Name = 'SearchKey'
    Expression = {
        $_.Item.Normalize([Text.NormalizationForm]::FormKC)
    }
}

Here, we retain the original Item and add a newSearchKeyis added. Since the original data is not lost, you can check later whether "they can be treated as the same string."

Note: Different original strings may result in the same SearchKey. Therefore, if uniqueness must be guaranteed by employee numbers or product codes, check whether duplicates occur after normalization before adopting them as keys.

Examples of data that must not be unified incorrectly

DataApproach to Applying NFKC
Product names for search assistanceNormalization can be tested in a column separate from the original string
Descriptions containing full-width alphanumeric charactersUnification can be convenient depending on the use case
Product codes and management numbersMust verify whether normalization causes distinct IDs to collide
Passwords and signature targetsDo not modify arbitrarily, as this changes meanings and verification results
File namesCheck for mismatches or collisions with existing paths

Furthermore, Unicode normalization does not unify "all visually similar characters." It is not a universal solution that resolves visually similar characters across different writing systems, invisible characters, or trailing spaces.

Failures and boundary conditions

Input is$null, calling string methods directly will fail. For data received from CSVs or APIs, always check for null first.

function ConvertTo-Nfkc([AllowNull()][string]$Value) {
    if ($null -eq $Value) { return $null }
    return $Value.Normalize([Text.NormalizationForm]::FormKC)
}

For invalid strings that fail Unicode normalization, consider exception handling when using them in production. Furthermore, PowerShell's-eqcomparison operators such as are case-insensitive by default, so if you need strict case checking, use-ceqor-cne.

Official Documentation

Document information

Article title
Standardizing Full-Width Alphanumeric Characters with PowerShell Unicode Normalization NFKC – A Non-Destructive Method for Source Data
Published
Updated
Source
https://papanda925.com/?p=18092&lang=en

License: Text and original figures for which this site holds the relevant rights are available under CC BY 4.0 , unless otherwise noted. This article may include content created or edited with generative AI. If code has a separate license notice or a linked GitHub repository license, that license takes precedence for the code. Quotations, third-party materials, images, and trademarks are excluded from this license. Usage policy

Copied title and URL