TY - RPRT TI - A unified analysis of value-function-based reinforcement- learning algorithms. AU - C Szepesvári AU - M L Littman PY - 1999 DO - 10.1162/089976699300016070 UR - https://pubmed.ncbi.nlm.nih.gov/10578043/ ID - 10578043 ER -