Quick recap, in case you're jumping in from Blog 4. Arthur Samuel's checkers program learned by guessing how good a board position was (value estimation). Donald Michie's matchbox machine, MENACE, learned by shifting which move it preferred (policy selection). Both worked. Both were fragile: Samuel needed hand-crafted features, and Michie needed... Weiterlesen: RL 5: Learning Automata and stochastic environments (1961–1974)
Intelligence View
⚡ tsecurity.de Intelligence
RL 5: Learning Automata and stochastic environments (1961–1974)
Quick recap, in case you're jumping in from Blog 4. Arthur Samuel's checkers program learned by guessing how good a board position was (value estimation).…