- mlx: match publisher tokenizer semantics
Honor pretokenizer stage order, split behavior, Unicode boundaries,
added-token normalization, and ranked BPE merges. Handle empty added
tokens and empty Metaspace input consistently.
Add shared Go/Python reference cases using published tokenizers, pulling
missing models directly and failing on errors, plus focused regressions
for configuration precedence, byte fallback, and parallel encoding.
- address comments