Câu 32: RAI: Risk and Artificial Intelligence
A financial firm is using a reinforcement learning trading model. The team determines that the reward value should be updated at the end of each scenario (episode). They test two methods: the first computes the reward using simple summation, and the second computes the reward using summation of discounted values. What…
Nội dung câu hỏi
A financial firm is using a reinforcement learning trading model. The team determines that the reward value should be updated at the end of each scenario (episode). They test two methods: the first computes the reward using simple summation, and the second computes the reward using summation of discounted values. What learning method(s) are used in this case?
Các lựa chọn
Đáp án được giữ gọn theo nhãn A, B, C, D trong phần bình chọn tương tác.
- A. Both methods are Monte Carlo. — đáp án hiện tại
- B. The first method is Temporal Difference, while the second one is Monte Carlo.
- C. The first method is Monte Carlo, while the second one is Temporal Difference.
- D. Both methods are Temporal Difference.
Cộng đồng
0 bình luận công khai. Tên thành viên được ẩn một phần.