Normalise text the way index terms are normalised, keeping a map back to the
original.
Needed because normalisation is not length-preserving — ü becomes ue, a
combining mark disappears, and İ lowercases to two code units — so an offset
found in the normalised string does not point at the same character in the
original. offsets[i] is the index in text of the character that produced
normalized[i], which is what lets a caller search normalised and then slice
the text the reader actually sees.
Normalise
textthe way index terms are normalised, keeping a map back to the original.Needed because normalisation is not length-preserving —
übecomesue, a combining mark disappears, andİlowercases to two code units — so an offset found in the normalised string does not point at the same character in the original.offsets[i]is the index intextof the character that producednormalized[i], which is what lets a caller search normalised and then slice the text the reader actually sees.