Most text processing assumes tokens are separated by spaces. Japanese and Korean quietly violate that assumption in different ways, and pipelines built on English defaults produce subtly wrong output rather than obvious errors.




Japanese: no spaces at all





私は本を読んでいます






No delimiters. Segmentation requires a morphological analyser...