Honestly it's not a new concept. this feature already existed in models before. problem was the models were just weak.

Looping only works if each attempt gets the agent closer to the correct solution. Earlier models weren't consistent enough for that. They often misunderstood feedback, repeated the same mistakes, or got stuck in an infinite loop....