tokenizer:token-string
(tokenizer:token-string tk id)
The token string the vocabulary gives id. For a byte-level BPE this is the byte-level spelling -- what the merge list is written in, not text: a space is Ġ and a newline Ċ, so that every byte is a printable, non-whitespace character. Use tokenizer:decode to turn ids back into text.