官术网_书友最值得收藏!

Identifying episodes

We mentioned earlier that the agent explores the environment in numerous trials-and-errors before it can learn to maximize its goals. Each such trial from start to finish is called an episode. The start location may or may not always be from the same location. Likewise, the finish or end of the episode can be a happy or sad ending.

A happy, or good, ending can be when the agent accomplishes its pre-defined goal, which could be successfully navigating to a final destination for a mobile robot, or successfully picking up a peg and placing it in a hole for an industrial robot arm, and so on. Episodes can also have a sad ending, where the agent crashes into obstacles or gets trapped in a maze, unable to get out of it, and so on.

In many RL problems, an upper bound in the form of a fixed number of time steps is generally specified for terminating an episode, although in others, no such bound exists and the episode can last for a very long time, ending with the accomplishment of a goal or by crashing into obstacles or falling off a cliff, or something similar. The Voyager spacecraft was launched by NASA in 1977, and has traveled outside our solar system – this is an example of a system with an infinite time episode.

We will next find out what a reward function is and why we need to discount future rewards. This reward function is the key, as it is the signal for the agent to learn.

主站蜘蛛池模板: 河东区| 祁东县| 曲阜市| 临湘市| 大丰市| 永春县| 西昌市| 阿图什市| 拜城县| 湖北省| 遵化市| 琼海市| 平湖市| 仁怀市| 宜君县| 定陶县| 宜川县| 周宁县| 邯郸市| 教育| 宜城市| 溆浦县| 涞水县| 兴海县| 石首市| 丹江口市| 元氏县| 绍兴市| 清流县| 碌曲县| 墨竹工卡县| 平邑县| 会东县| 威信县| 宝兴县| 夏津县| 和田市| 勐海县| 襄樊市| 杭州市| 宁河县|