Replace hardcoded CUDA-only graphsafe RNG operations with a device-generic
approach using a registry pattern. Device backends (TPU, etc.) can opt in
by calling register_graphsafe_rng_device_type() and providing
torch.<device_type>.default_generators with graphsafe_get/set_state().
Changes:
- Add registry (_GRAPHSAFE_RNG_DEVICE_TYPES) with CUDA registered by default
- Add helpers: get_default_generator(), get_generator_meta_val(),
supports_graphsafe_rng(), get_device_rng_state() - Add graphsafe_rng_device field to ViewAndMutationMeta schema
- Replace CUDA-specific code in partitioners.py, graph_compile.py,
runtime_wrappers.py with generic device-based equivalents - Register DispatchKey.PrivateUse1 for graphsafe_run_with_rng_state
- Add registry validation in BackendSelect dispatch