Make Inductor's accumulate_grad op take the current grad explicitly and return the updated grad, so functionalization and compiled autograd no longer rely on hidden Tensor.grad reads.
Remove the custom op output alias annotation and make the returned grad fresh: clone the current grad before accumulation, avoid aliasing new_grad on initialization,...