Totally model-free actor-critic recurrent neural-network reinforcement learning in non-Markovian domains