Least squares approximate policy iteration for learning bid prices in choice-based revenue management