Wrap triton_key in a cached function decorated with torch_key_cache.
This enables prefetching the triton key in a background thread via
triton_key.prefetch(), allowing it to be overlapped with other startup
work like torch_key computation.
Also moves torch_key_cache and _MISSING earlier in the file so they
can be used by triton_key before CacheBase.
Authored with Claude.
Pull Request resolved: #182188
Approved by: https://github.com/zou3519
ghstack dependencies: #182187