<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=AINAZHAJIMORADLOU</id>
	<title>UBC Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=AINAZHAJIMORADLOU"/>
	<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/Special:Contributions/AINAZHAJIMORADLOU"/>
	<updated>2026-09-16T14:07:18Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.9</generator>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/MarchApril2018&amp;diff=518892</id>
		<title>Course:CPSC522/MarchApril2018</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/MarchApril2018&amp;diff=518892"/>
		<updated>2018-04-27T18:29:56Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==March/April Assignment==&lt;br /&gt;
Your third and final assignment is to create a hypothesis related to the course and test it. Your result should describe the hypothesis and whether it works. The reader should be able to understand the background, what the the hypothesis is, whether the hypothesis is true, and the evidence you used to come to this conclusion. Your hypothesis can be theoretical or practical.&lt;br /&gt;
&lt;br /&gt;
==Extra Rules==&lt;br /&gt;
* You need to follow the rules on the main page and you should follow the guidelines there.&lt;br /&gt;
* Each page can have multiple authors (so it can be a group project).&lt;br /&gt;
* You need to add your page to the table of contents in a position that makes sense. Fell free to edit and change the structure of the table of content to give it a coherent structure.&lt;br /&gt;
* You will need to give a presentation of at most 4 minutes + 2 minutes for questions for each person (so if you have a group of 3, for example, you need a coherent presentation of 12 minutes where everyone participates); do not go over! If you would like to give a presentation during term time please contact David. For those who do not want to present during class time, we will have a presentation session during exam time.&lt;br /&gt;
* Please choose a topic that is different from other courses that you have done (or else you need to negotiate with the instructors to make sure you are not counting the same work multiple times). &lt;br /&gt;
* You should refer to wiki pages and to other research papers as appropriate. &lt;br /&gt;
===Key Dates===&lt;br /&gt;
* April 3 - last day to choose pages&lt;br /&gt;
* April 16 - First Draft ready for critiquing. Each page has a number Mn on the home page. &lt;br /&gt;
&lt;br /&gt;
- If a page M(n) has a single author, you will critique pages M(n-1), M(n+1) and M(n+6) where each addition is mod 15 (as there are 15 pages numbered M0 to M14). &lt;br /&gt;
&lt;br /&gt;
- If a page M(n) has multiple authors, you will between you critique pages M(n-1), M(n+1), M(n+6), M(n+8), M(n+10), M(n+13) all mod 15.  You decide between you who does which pages.&lt;br /&gt;
&lt;br /&gt;
Write your comments in the discussion tab of the page. Please give constructive feedback --- give the sort of feedback you would like to receive --- and answer the questions on the evaluation page.  Please feel free to respond to them there too, and actually have a discussion. The critiques are not meant to be anonymous; you are meant to be helping each other.&lt;br /&gt;
* April 17 10:00am-12:30 - Presentations&lt;br /&gt;
* April 19 - Critiques due&lt;br /&gt;
* April 23 - Final pages ready for peer marking&lt;br /&gt;
* April 28 - Peer marking completed. Use the template at http://cs.ubc.ca/~poole/cs522/2018/project_eval.py&lt;br /&gt;
&lt;br /&gt;
===Marking Scheme===&lt;br /&gt;
Here is a tentative marking scheme. This is subject to change. Feel free to add questions, and edit the questions if they do not make sense.&lt;br /&gt;
&lt;br /&gt;
On a scale of 1 to 5, where 1 means &amp;quot;strongly disagree&amp;quot; and 5 means &amp;quot;strongly agree&amp;quot; please rate and comment on the following:&lt;br /&gt;
* The topic is relevant for the course.&lt;br /&gt;
* The writing is clear and the English is good.&lt;br /&gt;
* The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds).&lt;br /&gt;
* The formalism (definitions, mathematics) was well chosen to make the page easier to understand.&lt;br /&gt;
* The abstract is a concise and clear summary.&lt;br /&gt;
* There were appropriate (original) examples that helped make the topic clear.&lt;br /&gt;
* There was appropriate use of (pseudo-) code.&lt;br /&gt;
* It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic).&lt;br /&gt;
* It is correct.&lt;br /&gt;
* It was neither too short nor too long for the topic.&lt;br /&gt;
* It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page).&lt;br /&gt;
* It links to appropriate other pages in the wiki.&lt;br /&gt;
* The references and links to external pages are well chosen.&lt;br /&gt;
* I would recommend this page to someone who wanted to find out about the topic.&lt;br /&gt;
* This page should be highlighted as an exemplary page for others to emulate.&lt;br /&gt;
&lt;br /&gt;
If I was grading it out of 20, I would give it:&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/SLAM_And_Sensor_Quality/Critique_(2)&amp;diff=518466</id>
		<title>Thread:Course talk:CPSC522/SLAM And Sensor Quality/Critique (2)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/SLAM_And_Sensor_Quality/Critique_(2)&amp;diff=518466"/>
		<updated>2018-04-23T06:56:53Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: New thread: Critique&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Overall, It&#039;s a good written page and the topic is interesting. However, I have some comments that might be useful:&lt;br /&gt;
&lt;br /&gt;
1. There are some typos and grammatical errors in the page (e.g. based off, an simulated robot, etc.).&lt;br /&gt;
&lt;br /&gt;
2. Although the problem is clear, it&#039;s still not clear why decreasing noise like nearly perfect sensors result in worse localization which then leads to robot being lost.&lt;br /&gt;
&lt;br /&gt;
3. It would&#039;ve been a good idea to write a conclusion/discussion section that summarizes the mentioned experiments and states the final conclusion on such systems.&lt;br /&gt;
&lt;br /&gt;
---------------------------------------------------------------------------------------------------&lt;br /&gt;
Marking scheme:&lt;br /&gt;
 * The topic is relevant for the course. 5&lt;br /&gt;
 * The writing is clear and the English is good. 4&lt;br /&gt;
 * The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5&lt;br /&gt;
 * The formalism (definitions, mathematics) was well chosen to make the page easier to understand. N/A&lt;br /&gt;
 * The abstract is a concise and clear summary. 4&lt;br /&gt;
 * There were appropriate (original) examples that helped make the topic clear. 5&lt;br /&gt;
 * There was appropriate use of (pseudo-) code. 2&lt;br /&gt;
 * It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). N/A&lt;br /&gt;
 * It is correct. 5 &lt;br /&gt;
 * It was neither too short nor too long for the topic. 5&lt;br /&gt;
 * It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page). 5&lt;br /&gt;
 * It links to appropriate other pages in the wiki. 3&lt;br /&gt;
 * The references and links to external pages are well chosen. 4&lt;br /&gt;
 * I would recommend this page to someone who wanted to find out about the topic. 4&lt;br /&gt;
 * This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
 If I was grading it out of 20, I would give it: 17&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Experiments_with_Reinforcement_Learning/Critique_3&amp;diff=518461</id>
		<title>Thread:Course talk:CPSC522/Experiments with Reinforcement Learning/Critique 3</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Experiments_with_Reinforcement_Learning/Critique_3&amp;diff=518461"/>
		<updated>2018-04-23T04:02:07Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: New thread: Critique 3&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The page looks good! It has a good coverage and the explanations are thorough. I didn&#039;t actually find anything to critique except the fact that the page is incomplete and you can talk more about the results of the things you tried. Maybe talking a little more about frameworks/implementation details or challenges would be useful.&lt;br /&gt;
&lt;br /&gt;
---------------------------------------------------------------------------------------------------&lt;br /&gt;
Marking scheme:&lt;br /&gt;
 * The topic is relevant for the course. 5&lt;br /&gt;
 * The writing is clear and the English is good. 5&lt;br /&gt;
 * The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5&lt;br /&gt;
 * The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4&lt;br /&gt;
 * The abstract is a concise and clear summary. 4&lt;br /&gt;
 * There were appropriate (original) examples that helped make the topic clear. 4&lt;br /&gt;
 * There was appropriate use of (pseudo-) code. N/A&lt;br /&gt;
 * It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). 3&lt;br /&gt;
 * It is correct. 5 &lt;br /&gt;
 * It was neither too short nor too long for the topic. 5&lt;br /&gt;
 * It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page). 5&lt;br /&gt;
 * It links to appropriate other pages in the wiki. 2&lt;br /&gt;
 * The references and links to external pages are well chosen. N/A&lt;br /&gt;
 * I would recommend this page to someone who wanted to find out about the topic. 4&lt;br /&gt;
 * This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
 If I was grading it out of 20, I would give it: 18&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518460</id>
		<title>Thread:Course talk:CPSC522/An evaluation on selecting and applying Recommendation Methods/Critique</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518460"/>
		<updated>2018-04-23T04:01:27Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Hi guys, &lt;br /&gt;
&lt;br /&gt;
Overal, it&#039;s a thorough survey on different methods of filtering. I have some comments that may be useful to you though:&lt;br /&gt;
&lt;br /&gt;
1. There are some grammatical errors/typos such as &amp;quot;based off&amp;quot;, &amp;quot;Be base our implementation based off the description of both methods&amp;quot;, etc.&lt;br /&gt;
&lt;br /&gt;
2. You might want to write references in [1] format instead of superscript in the original context. (e.g. in [1], they focus on ...)&lt;br /&gt;
&lt;br /&gt;
3. It would be nice to have some quantitative results as well as qualitative results like using tables to report accuracy on some subset of dataset for all proposed algorithms and comparing them.&lt;br /&gt;
&lt;br /&gt;
---------------------------------------------------------------------------------------------------&lt;br /&gt;
Marking scheme:&lt;br /&gt;
 * The topic is relevant for the course. 5&lt;br /&gt;
 * The writing is clear and the English is good. 4&lt;br /&gt;
 * The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5&lt;br /&gt;
 * The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4&lt;br /&gt;
 * The abstract is a concise and clear summary. 4&lt;br /&gt;
 * There were appropriate (original) examples that helped make the topic clear. 4&lt;br /&gt;
 * There was appropriate use of (pseudo-) code. 3&lt;br /&gt;
 * It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). N/A&lt;br /&gt;
 * It is correct. 5 &lt;br /&gt;
 * It was neither too short nor too long for the topic. 5&lt;br /&gt;
 * It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page). 5&lt;br /&gt;
 * It links to appropriate other pages in the wiki. 3&lt;br /&gt;
 * The references and links to external pages are well chosen. 4&lt;br /&gt;
 * I would recommend this page to someone who wanted to find out about the topic. 4&lt;br /&gt;
 * This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
 If I was grading it out of 20, I would give it: 17&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518455</id>
		<title>Thread:Course talk:CPSC522/An evaluation on selecting and applying Recommendation Methods/Critique</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518455"/>
		<updated>2018-04-23T03:45:20Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Hi guys, &lt;br /&gt;
&lt;br /&gt;
Overal, it&#039;s a thorough survey on different methods of filtering. I have some comments that may be useful to you though:&lt;br /&gt;
&lt;br /&gt;
1. There are some grammatical errors/typos such as &amp;quot;based off&amp;quot;, &amp;quot;Be base our implementation based off the description of both methods&amp;quot;, etc.&lt;br /&gt;
&lt;br /&gt;
2. You might want to write references in [1] format instead of superscript in the original context. (e.g. in [1], they focus on ...)&lt;br /&gt;
&lt;br /&gt;
3. It would be nice to have some quantitative results as well as qualitative results like using tables to report accuracy on some subset of dataset for all proposed algorithms and comparing them.&lt;br /&gt;
&lt;br /&gt;
---------------------------------------------------------------------------------------------------&lt;br /&gt;
Marking scheme:&lt;br /&gt;
 * The topic is relevant for the course. 5&lt;br /&gt;
 * The writing is clear and the English is good. 4&lt;br /&gt;
 * The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5&lt;br /&gt;
 * The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4&lt;br /&gt;
 * The abstract is a concise and clear summary. 4&lt;br /&gt;
 * There were appropriate (original) examples that helped make the topic clear. 4&lt;br /&gt;
 * There was appropriate use of (pseudo-) code. 3&lt;br /&gt;
 * It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). N/A&lt;br /&gt;
 * It is correct. 5 &lt;br /&gt;
 * It was neither too short nor too long for the topic. 3 (It&#039;s a little short so far, but it seems obvious you intend to expand it a bit before the final version)&lt;br /&gt;
 * It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page). 5&lt;br /&gt;
 * It links to appropriate other pages in the wiki. N/A&lt;br /&gt;
 * The references and links to external pages are well chosen. N/A&lt;br /&gt;
 * I would recommend this page to someone who wanted to find out about the topic. 4&lt;br /&gt;
 * This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
 If I was grading it out of 20, I would give it: 17&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518454</id>
		<title>Thread:Course talk:CPSC522/An evaluation on selecting and applying Recommendation Methods/Critique</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods/Critique&amp;diff=518454"/>
		<updated>2018-04-23T03:44:54Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: New thread: Critique&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Hi guys, &lt;br /&gt;
Overal, it&#039;s a thorough survey on different methods of filtering. I have some comments that may be useful to you though:&lt;br /&gt;
1. There are some grammatical errors/typos such as &amp;quot;based off&amp;quot;, &amp;quot;Be base our implementation based off the description of both methods&amp;quot;, etc.&lt;br /&gt;
2. You might want to write references in [1] format instead of superscript in the original context. (e.g. in [1], they focus on ...)&lt;br /&gt;
3. It would be nice to have some quantitative results as well as qualitative results like using tables to report accuracy on some subset of dataset for all proposed algorithms and comparing them.&lt;br /&gt;
---------------------------------------------------------------------------------------------------&lt;br /&gt;
Marking scheme:&lt;br /&gt;
 * The topic is relevant for the course. 5&lt;br /&gt;
 * The writing is clear and the English is good. 4&lt;br /&gt;
 * The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5&lt;br /&gt;
 * The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4&lt;br /&gt;
 * The abstract is a concise and clear summary. 4&lt;br /&gt;
 * There were appropriate (original) examples that helped make the topic clear. 4&lt;br /&gt;
 * There was appropriate use of (pseudo-) code. 3&lt;br /&gt;
 * It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). N/A&lt;br /&gt;
 * It is correct. 5 &lt;br /&gt;
 * It was neither too short nor too long for the topic. 3 (It&#039;s a little short so far, but it seems obvious you intend to expand it a bit before the final version)&lt;br /&gt;
 * It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page). 5&lt;br /&gt;
 * It links to appropriate other pages in the wiki. N/A&lt;br /&gt;
 * The references and links to external pages are well chosen. N/A&lt;br /&gt;
 * I would recommend this page to someone who wanted to find out about the topic. 4&lt;br /&gt;
 * This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
 If I was grading it out of 20, I would give it: 17&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518449</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518449"/>
		<updated>2018-04-23T03:27:37Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Approach */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
[[File:Motion_demo.gif|300px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity and the spent time, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Demo motion2.gif|300px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518448</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518448"/>
		<updated>2018-04-23T03:26:15Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Simulated Environment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
[[File:Motion_demo.gif|300px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Demo motion2.gif|300px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518146</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518146"/>
		<updated>2018-04-19T08:21:43Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
[[File:Motion_demo.gif|300px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Demo motion2.gif|300px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518145</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518145"/>
		<updated>2018-04-19T08:17:47Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Future Work */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|280px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|280px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|280px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|100px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
&lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518144</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518144"/>
		<updated>2018-04-19T08:16:37Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|280px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|280px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|280px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
&lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518143</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518143"/>
		<updated>2018-04-19T08:15:01Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: Undo revision 518142 by AINAZHAJIMORADLOU (talk)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|250px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|250px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
&lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518142</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518142"/>
		<updated>2018-04-19T08:13:47Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|250px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|250px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518140</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518140"/>
		<updated>2018-04-19T08:12:19Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|250px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|250px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
Figure 4 shows an example of the results obtained by the baseline. In this case, the velocities along each axis are increased to the same extent. Therefore, the robot moves in the y = x direction. However, this is not optimal as the robot&#039;s direction changes every now and then. Depending on the direction, the velocities should be amplified accordingly. To solve this, we can simply find the angle between current position of the robot and the desired direction which is the goal. Cosine and Sine of this angle will be the multipliers of the velocities in x and y axes. The updated results are shown in figure 5. The direction of the user in this figure is at point (1, 0). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|200px|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
&lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles. &lt;br /&gt;
&lt;br /&gt;
Besides, speaking in terms of model, we can smooth the obtained parameters of the state by using moving average instead of finding a desired velocity in the specified margin. To further extend the project we should focus on adapting the model to more complex simulation environments as the one shown in figure 6.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518123</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518123"/>
		<updated>2018-04-19T05:26:07Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|300px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|300px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]] [[File:Demo motion2.gif|300px|thumb|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518122</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518122"/>
		<updated>2018-04-19T05:22:45Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: AINAZHAJIMORADLOU uploaded a new version of &amp;amp;quot;File:Motion demo.gif&amp;amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518121</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518121"/>
		<updated>2018-04-19T05:20:56Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|380px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|380px|thumb|left|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
[[File:Demo motion2.gif|380px|thumb|left|&#039;&#039;Figure 5.&#039;&#039; Simple robot motion with different velocities along each axis.]]&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Demo_motion2.gif&amp;diff=518120</id>
		<title>File:Demo motion2.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Demo_motion2.gif&amp;diff=518120"/>
		<updated>2018-04-19T05:20:08Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Simple robot motion with different velocities along each axis.}}&lt;br /&gt;
|date=2018-04-18 22:19:11&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518119</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518119"/>
		<updated>2018-04-19T05:16:22Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|380px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|380px|thumb|&#039;&#039;Figure 4.&#039;&#039; A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518118</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518118"/>
		<updated>2018-04-19T05:15:30Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: AINAZHAJIMORADLOU uploaded a new version of &amp;amp;quot;File:Motion demo.gif&amp;amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518117</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518117"/>
		<updated>2018-04-19T05:14:32Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: AINAZHAJIMORADLOU uploaded a new version of &amp;amp;quot;File:Motion demo.gif&amp;amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518116</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518116"/>
		<updated>2018-04-19T05:13:46Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|380px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; Magnitude of velocity of the user at each time step.]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
[[File:Motion_demo.gif|thumb|A simple robot motion based on the proposed baseline.]]&lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518115</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518115"/>
		<updated>2018-04-19T05:10:22Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: AINAZHAJIMORADLOU uploaded a new version of &amp;amp;quot;File:Motion demo.gif&amp;amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518114</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518114"/>
		<updated>2018-04-19T05:07:56Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: AINAZHAJIMORADLOU uploaded a new version of &amp;amp;quot;File:Motion demo.gif&amp;amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518113</id>
		<title>File:Motion demo.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_demo.gif&amp;diff=518113"/>
		<updated>2018-04-19T05:03:06Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=A simple robot motion based on the proposed baseline.}}&lt;br /&gt;
|date=2018-04-18 22:02:34&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518110</id>
		<title>Course:CPSC522/Learning User Preferences of Motion Control</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Learning_User_Preferences_of_Motion_Control&amp;diff=518110"/>
		<updated>2018-04-19T04:35:18Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Title ==&lt;br /&gt;
The goal of this project is to add intervention to a power wheelchair system so that when the user expresses frustration due to limited joystick control abilities, the system can adjust the control to match the users preference.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou, Jocelyn Minns&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
[[File:PWCPartsDiagram.jpg|250px||thumb|&#039;&#039;Figure 1.&#039;&#039; Design of a Power Wheelchair. Image source from usmedicalsupplies.com]]&lt;br /&gt;
&lt;br /&gt;
Modern power wheelchairs are an essential tool to many who are mobility challenged. However, while this gives freedom to some who would otherwise be restricted in their day to day life, other users are still limited from safely operating these large, heavy and powerful machines. Users that experience motor and/or cognitive impairments may be be unable to operate a power wheelchair to the accuracy that is required. &lt;br /&gt;
Most common power wheelchairs are operated using a joystick controller to steer the motors. If the user is unable to perform the small movements for persuasion control and the large movements to achieve reasonable speed, then they could become frustrated and may compromise their own safety. The goal of this project is to add intervention to the system so that when the user expresses frustration with the control of the power wheelchair, the system can adjust to match the users preference.&lt;br /&gt;
&lt;br /&gt;
=== Simulated Environment ===  &lt;br /&gt;
[[File:Test-drive-holonomic.gif|380px||thumb|left|&#039;&#039;Figure 2.&#039;&#039; A simulated holonomic robot exploring the environment.]]&lt;br /&gt;
&lt;br /&gt;
The most common form of a power wheelchair can be modelled as a differential drive robot. Meaning that the robot has two main wheels attached to the sides of the robot each with its own motor. A wheelchair also contains four caster wheels at the front and back to ensure stability and balance. The caster wheels generally do not effect the motion of the chair since they can rotate freely and are not attached to any motor. Since each of the two main wheel are attached to its own motor, they can rotate in opposite directions or at different velocities allowing the user to turn or even spin in place. This model proposes its own set of challenges in simulation. Accurately mapping the linear and rotational velocities given the user input from the joystick control can be a large problem on its own.&lt;br /&gt;
&lt;br /&gt;
To test our initial approach, it is desirable to start with a simpler model than the one described above. Using a holonomic robot design lifts some of the resections of the differential drive design allowing us to easily simulate the movement of the robot. With this design, the robot is free to move in any direction in the x-y plane. We can then focus on the linear velocity only and adapt to the users preferences accordingly.&lt;br /&gt;
&lt;br /&gt;
For our simulation, we are not able to use real power wheelchair users to provide feedback for our system. Instead we simulate users such that each user will have a set level of ability to use the joystick controls and a desired speed that they wish to go. The desired speed remains unknown to the system, the only feedback that the user gives is a verbal description of the speed as as too fast or too slow. The joystick controls are fed to the system at 30 Hz, however the user is unable to provide feedback at that rate so there is some delay to the users commands based on the accurate movement of the chair.&lt;br /&gt;
&lt;br /&gt;
=== Approach === &lt;br /&gt;
&lt;br /&gt;
[[File:Motion control plot.png|380px||thumb|&#039;&#039;Figure 3.&#039;&#039; motion control plot]]&lt;br /&gt;
The first models we considered were Markov models such as [https://en.wikipedia.org/wiki/Markov_chain Markov chains], [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and Continuous Markov process. Since, obtaining the parameters of each state such as velocity and position depends on previous states. However, we encountered a problem when fitting data with these models due to the nature of the proposed problem. We have to get the feedback from the user in form of verbal complaint, but these models do not support having such inputs. Therefore, we decided to go with a simple version of a baseline model.&lt;br /&gt;
&lt;br /&gt;
Each state in the model specifies current velocity, position and the direction the user is going with joystick controllers. A desired linear velocity exists which is specified by user as the input to the model. So, based on this linear speed, feedbacks are divided into three categories: fast, good and slow. The model&#039;s approach is to start with increasing the velocity until the feedback from the user states that the speed is too fast. The point where this feedback is achieved, high margin of the speed is obtained. Low margin is the previous state&#039;s velocity. Based on these margins, the desired speed of the user is obtained by a simple binary search. When the desired speed is achieved, user will give a feedback (&amp;quot;good speed&amp;quot;) and the velocity will remain constant until the user&#039;s direction changes. So at each time step, based on current velocity, the position of the user is updated in the plot. This process will repeat until the user stops or reaches the desired speed. The magnitude of velocity at each time step is plotted in figure 3.&lt;br /&gt;
&lt;br /&gt;
=== Results === &lt;br /&gt;
&lt;br /&gt;
=== Future Work === &lt;br /&gt;
[[File:Test-drive-nonholonomic.gif|thumb|left|A simulated differential drive robot using the RobotPy package.]]&lt;br /&gt;
&lt;br /&gt;
The two main focuses for future work is to improve the robot design of the power wheelchair and to improve the simulated user. &lt;br /&gt;
Using a predesigned robot simulator such as RobotPy, we are able to simulate a differential drive robot. However, the mapping of the joystick controller to the movement is highly inaccurate to a real wheelchair system. Power wheelchairs inherently apply some mapping function from the control to the motors to provide the user with safer control. Examples of this is the decreased velocity when backing up or the decreased rotational velocity when spinning in place. To build a mapping of the control to the motors is a non trivial job, but would be necessary to use a more accurate model. Once the model was in place, we would be able to adjust the control to accommodate the users preference for both the linear and rotational velocities. &lt;br /&gt;
To improve the simulated user, we would add noise to the users inputs/outputs. The user perception of the velocity would have some variation as well as the users ability to use the joystick control. The user would also want to adjust their desired speed given different environments or different obstacles.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Motion_control_plot.png&amp;diff=518107</id>
		<title>File:Motion control plot.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Motion_control_plot.png&amp;diff=518107"/>
		<updated>2018-04-19T04:22:01Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=motion control plot}}&lt;br /&gt;
|date=2018-04-18 21:21:25&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:AINAZHAJIMORADLOU|AINAZHAJIMORADLOU]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-3.0}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/StudentPresentations2018&amp;diff=517713</id>
		<title>Course:CPSC522/StudentPresentations2018</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/StudentPresentations2018&amp;diff=517713"/>
		<updated>2018-04-17T04:09:38Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==CPSC 522 - Project Presentation Schedule 2018-04-17 ==&lt;br /&gt;
The project presentations will be on 2018-04-17 at 10:00 in room 146. Please add you name either before or after the break. You are only allowed to add your name to one of the times marked with #&lt;br /&gt;
# Borna &lt;br /&gt;
# Vanessa&lt;br /&gt;
# Fabian&lt;br /&gt;
# Wenyi&lt;br /&gt;
# May&lt;br /&gt;
# David&lt;br /&gt;
# Ekta&lt;br /&gt;
# Jocelyn, Ainaz&lt;br /&gt;
# &lt;br /&gt;
# Break&lt;br /&gt;
# Surbhi&lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
# &lt;br /&gt;
#&lt;br /&gt;
&lt;br /&gt;
==CPSC 522 - January to April 2018 Student Presentation Schedule==&lt;br /&gt;
Please sign up for a date you want to present. There is a maximum of three students per slot. (Hopefully most slots will end up with three presenters).  The three students &#039;&#039;&#039;must coordinate&#039;&#039;&#039; to present the two papers in 60 minutes.  All three students must present.  The two papers are related. It is suggested that one student gives 20 minutes background, and the other students each present 20 minutes for one of the papers. But the format is up to the presenters.&lt;br /&gt;
&lt;br /&gt;
The schedule of which papers are presented each day is at  http://www.cs.ubc.ca/~poole/cs522/2018/schedule.html&lt;br /&gt;
&lt;br /&gt;
==Dates==&lt;br /&gt;
* Jan 16 - Everyone!&lt;br /&gt;
* Jan 23 -&lt;br /&gt;
* Jan 30 -&lt;br /&gt;
* Feb 8 - David Johnson, May Young&lt;br /&gt;
* Feb 15 - Kevin Dsouza, Bronson Bouchard&lt;br /&gt;
* Mar 1 - Kumseok Jung, Julin Song, Ekta Aggarwal&lt;br /&gt;
* Mar 8 - Fabian Ruffy, Alistair Wick, Gudbrand Tandberg&lt;br /&gt;
* Mar 15 - Borna Ghotbi, Vanessa Putnam&lt;br /&gt;
* Mar 22 - Amin Aghaee, Carl Kwan, Surbhi Palande&lt;br /&gt;
* Mar 29 - &lt;br /&gt;
* Apr 3 - Wenyi Wang, Ainaz Hajimoradlou, Jocelyn Minns&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506758</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506758"/>
		<updated>2018-03-23T20:25:50Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Abstract */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning models such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are part of deep reinforcement learning. In other words, deep reinforcement learning is a combination of reinforcement learning and deep learning techniques.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This page gives a short review of two papers in deep reinforcement learning. The first paper discusses a new approach for integrating deep learning techniques into reinforcement learning problems. The model is experimented on 7 different Atari games and is able to achieve higher scores in comparison to state of the art models at that time. The second paper expands this approach by 6 different techniques such as different architectures and sampling techniques on 57 Atari games. The model is able to achieve better scores compared to the previous approach.&lt;br /&gt;
===Builds on===&lt;br /&gt;
This page builds on general concepts such as [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] and [[Course:CPSC522/Neural_Network|Neural Networks]]. The [[Course:CPSC522/Convolutional_Neural_Networks|convolutional neural networks]] along with [[Course:CPSC522/Reinforcement_Learning_with_Function_approximation|function approximators]] are used to model games like [[Course:CPSC522/Deep_Neural_Network|Alpha go]] using reinforcement learning techniques.&lt;br /&gt;
&lt;br /&gt;
===Related Pages===&lt;br /&gt;
Games like [[Course:CPSC522/Deep_Neural_Network|Alpha go]] are some examples of applications of deep reinforcement learning. These models are usually based on non-linear [[Course:CPSC522/Reinforcement_Learning_with_Function_approximation|function approximators]] such as [[Course:CPSC522/Convolutional_Neural_Networks|convolutional neural networks]]. Moreover, the essence of deep reinforcement learning algorithms is dependent on  [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506757</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506757"/>
		<updated>2018-03-23T20:19:56Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Deep Reinforcement Learning */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning models such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are part of deep reinforcement learning. In other words, deep reinforcement learning is a combination of reinforcement learning and deep learning techniques.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This page gives a short review of two papers in deep reinforcement learning. The first paper discusses a new approach for integrating deep learning techniques into reinforcement learning problems. The model is experimented on 7 different Atari games and is able to achieve higher scores in comparison to state of the art models at that time. The second paper expands this approach by 6 different techniques such as different architectures and sampling techniques on 57 Atari games. The model is able to achieve better scores compared to the previous approach.&lt;br /&gt;
===Builds on===&lt;br /&gt;
This page builds on general concepts such as [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] and [[Course:CPSC522/Neural_Network|Neural Networks]]. The [[Course:CPSC522/Convolutional_Neural_Networks|convolutional neural networks]] along with [[Course:CPSC522/Reinforcement_Learning_with_Function_approximation|function approximators]] are used to model games like [[Course:CPSC522/Deep_Neural_Network|Alpha go]] using reinforcement learning techniques.&lt;br /&gt;
&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506111</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506111"/>
		<updated>2018-03-20T19:35:58Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Builds on */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning techniques such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are called deep reinforcement learning.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This page gives a short review of two papers in deep reinforcement learning. The first paper discusses a new approach for integrating deep learning techniques into reinforcement learning problems. The model is experimented on 7 different Atari games and is able to achieve higher scores in comparison to state of the art models at that time. The second paper expands this approach by 6 different techniques such as different architectures and sampling techniques on 57 Atari games. The model is able to achieve better scores compared to the previous approach.&lt;br /&gt;
===Builds on===&lt;br /&gt;
This page builds on general concepts such as [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] and [[Course:CPSC522/Neural_Network|Neural Networks]]. The [[Course:CPSC522/Convolutional_Neural_Networks|convolutional neural networks]] along with [[Course:CPSC522/Reinforcement_Learning_with_Function_approximation|function approximators]] are used to model games like [[Course:CPSC522/Deep_Neural_Network|Alpha go]] using reinforcement learning techniques.&lt;br /&gt;
&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506110</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=506110"/>
		<updated>2018-03-20T19:35:37Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Abstract */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning techniques such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are called deep reinforcement learning.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This page gives a short review of two papers in deep reinforcement learning. The first paper discusses a new approach for integrating deep learning techniques into reinforcement learning problems. The model is experimented on 7 different Atari games and is able to achieve higher scores in comparison to state of the art models at that time. The second paper expands this approach by 6 different techniques such as different architectures and sampling techniques on 57 Atari games. The model is able to achieve better scores compared to the previous approach.&lt;br /&gt;
===Builds on===&lt;br /&gt;
This page builds on general concepts such as [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] and [[Course:CPSC522/Neural_Network|Neural Networks]]. The [[Course:CPSC522/Convolutional_Neural_Networks|convolutional neural networks]] along with [[Course:CPSC522/Reinforcement_Learning_with_Function_approximation|function approximators]] are used to model games like [[Course:CPSC522/Deep_Neural_Network|Alpha go]] using reinforcement learning techniques.&lt;br /&gt;
,===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505881</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505881"/>
		<updated>2018-03-19T21:26:38Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Abstract */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning techniques such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are called deep reinforcement learning.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This page gives a short review of two papers in deep reinforcement learning. The first paper discusses a new approach for integrating deep learning techniques into reinforcement learning problems. The model is experimented on 7 different Atari games and is able to achieve higher scores in comparison to state of the art models at that time. The second paper expands this approach by 6 different techniques such as different architectures and sampling techniques on 57 Atari games. The model is able to achieve better scores compared to the previous approach.&lt;br /&gt;
===Builds on===&lt;br /&gt;
This page builds on general concepts of deep learning, reinforcement learning and neural network.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505880</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505880"/>
		<updated>2018-03-19T21:17:28Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Deep Reinforcement Learning */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
Deep learning techniques such as [https://en.wikipedia.org/wiki/Artificial_neural_network neural networks] that are used in [https://en.wikipedia.org/wiki/Reinforcement_learning reinforcement learning] to learn control policies directly from high dimensional data are called deep reinforcement learning.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505878</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505878"/>
		<updated>2018-03-19T21:06:38Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* To Add */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505877</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505877"/>
		<updated>2018-03-19T21:06:19Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Annotated Bibliography */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505876</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505876"/>
		<updated>2018-03-19T21:05:59Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Discussion */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505875</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505875"/>
		<updated>2018-03-19T21:05:31Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. &lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
=== Discussion === &lt;br /&gt;
The first paper introduces a new approach to deep reinforcement learning by a non-linear function approximator. Although, the chances of getting a deep network to work with reinforcement learning algorithms is very low, the paper was able to achieve not only reasonably good results but higher scores compared to the state of the art approaches which was a big breakthrough in this area. On the other hand, the second paper is actually just an integration method of different existing extensions to the proposed algorithm in the first paper. Each of these extensions have been used to improve one of the limitations of deep Q-learning. Therefore, having a model composed of all these extensions has resulted in improvement of the original model.&lt;br /&gt;
&lt;br /&gt;
== Discussion ==&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505868</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505868"/>
		<updated>2018-03-19T20:47:10Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Paper I, Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Paper ||, Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
&lt;br /&gt;
== Discussion ==&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505867</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505867"/>
		<updated>2018-03-19T20:44:51Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
[[File:Rainbow DQN.png|1200px||thumb|&#039;&#039;&#039;Figure 2.&#039;&#039;&#039; First row shows the number of games where each agent was able to achieve at least a fraction of human performance as a function of time. The bottom row compares performance of rainbow to the models where some component is removed from the original model&amp;lt;ref name=&amp;quot;paper2&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;.]]&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Rainbow_DQN.png&amp;diff=505864</id>
		<title>File:Rainbow DQN.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Rainbow_DQN.png&amp;diff=505864"/>
		<updated>2018-03-19T20:38:01Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Deep Q network}}&lt;br /&gt;
|date=2018-03-19 13:36:18&lt;br /&gt;
|source=Rainbow: Combining Improvements in Deep Reinforcement Learning&lt;br /&gt;
|author=M. Hessel, et. al.&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{subst:uwl}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505863</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505863"/>
		<updated>2018-03-19T20:36:31Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Rainbow: Combining Improvements in Deep Reinforcement Learning */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning: DQN accumulates a single reward and then uses greedy action at the next step to bootstrap. An an alternative, forward-view multi-step targets&amp;lt;ref&amp;gt;Richard S Sutton, Learning to predict by the methods of temporal differences&amp;lt;/ref&amp;gt; can be used for optimizing the parameters which can lead to faster learning with suitable &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Therefore, the reward function will be defined as follows. &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
R_t^{(n)} \equiv \sum_{k=0}^{n-1} \gamma_t^{(k)}R_{t+k+1}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Distributional RL: Instead of approximating the expected return, the distribution of returns can be estimated&amp;lt;ref&amp;gt;Bellemare, Dabney, and Munos, Distributional Reinforcement Learning with Quantile Regression&amp;lt;/ref&amp;gt;. Therefore, the output of the network would be the distribution over each action.&lt;br /&gt;
* Noisy Nets: Exploration using &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approaches is not suitable for situations where a lot of actions should be executed to collect the first reward. In these cases a model with a linear layer composed of a deterministic and noisy stream is proposed&amp;lt;ref&amp;gt;Fortunato et. al., Noisy Networks for Exploration&amp;lt;/ref&amp;gt;. The model makes the network learn that it can ignore the noisy stream. The extent varies as the state space changes. This will allow state-conditional exploration with a form of self annealing. &amp;lt;math&amp;gt;\epsilon^b&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\epsilon^w&amp;lt;/math&amp;gt; are random variables and &amp;lt;math&amp;gt;\odot&amp;lt;/math&amp;gt; is the element-wise multiplication.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y = (b+Wx) + (b_{noisy}\odot\epsilon^b + (W_{noisy}\odot \epsilon^w)x)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The paper uses the integrated version of all above extensions as the original model. They combine a multi-step distributional loss with double Q-learning to choose the best action. They use prioritized experience replay to store the most frequent transitions in the memory. The network architecture is a dueling network which is adapted to use return distributions and all the linear layers are replaced with their equivalent noisy nets.&lt;br /&gt;
====Experiments====&lt;br /&gt;
The model is evaluated on 57 Atari games. The results are compared to human performance in all games. As the model is actually a combination of 6 other models, the number of hyper-parameters is large. Therefore, some tuning is needed for each component. The initial values are the values proposed by the original papers corresponding to each component. Then, the most sensitive ones are tuned using manual [https://en.wikipedia.org/wiki/Coordinate_descent coordinate descent]. The hyper parameters are the same across all games. Figure 2, top row, shows different plots for several agents, each plot represents the number of games the agent was able to achieve a given fraction of human performance. Bottom row shows the significance of each component in the integrated model by removing each component from the model and observing the performance.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505835</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505835"/>
		<updated>2018-03-19T19:10:06Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right]. \quad(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].\quad(2)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right] \quad(3)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right] \quad(4)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
[[File:Qlearning table.png|700px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
=== Rainbow: Combining Improvements in Deep Reinforcement Learning ===&lt;br /&gt;
The main essence of this paper is based on an integrated method for extending the DQN algorithm, proposed in the previous paper. Generally, at each time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; the environment provides agent with an observation &amp;lt;math&amp;gt;S_t&amp;lt;/math&amp;gt;. In response, the agent selects an action &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; which provides the next reward &amp;lt;math&amp;gt;R_{t+1}&amp;lt;/math&amp;gt;, discount factor &amp;lt;math&amp;gt;\gamma_{t+1}&amp;lt;/math&amp;gt; and state &amp;lt;math&amp;gt;S_{t+1}&amp;lt;/math&amp;gt;. This can be formalized as a [https://en.wikipedia.org/wiki/Markov_decision_process Markov Decision process] &amp;lt;math&amp;gt;&amp;lt;S, A, T, r, \gamma&amp;gt;&amp;lt;/math&amp;gt;. The members of the tuple correspond to the finite set of states, finite set of actions, transition function, reward function and the discount factor respectively. The action is chosen based on the policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which is the distribution over actions. In DQN, parameters are optimized using the loss function proposed in equation (3). There are several limitations to this algorithm which are addressed by different methods. Some of these extensions are as follows:&lt;br /&gt;
* Double Q-learning: Learning in DQN is affected by an overestimation bias in the maximization step performed on loss function in equation (3). This can be addressed by decoupling the selection of the action from its evaluation. Therefore, the target value in the loss function will be changed as follows &amp;lt;ref&amp;gt;Hado van Hasselt, Arthur Guez, David Silver, Deep Reinforcement Learning with Double Q-learning&amp;lt;/ref&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma Q(s&#039;, \arg\max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}); \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
* Prioritized replay: In DQN, the samples are chosen uniformly from the replay memory. However, the samples from which the network learns the most should be chosen more frequently. The idea is using prioritized experience replay&amp;lt;ref&amp;gt;Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver, Prioritized Experience Replay&amp;lt;/ref&amp;gt;. Therefore, transitions that are relevant to the last encountered [https://en.wikipedia.org/wiki/Temporal_difference_learning TD] error will have higher probabilities and are chosen most of the times. &lt;br /&gt;
* Dueling networks: This is a neural network architecture which is designed for value based reinforcement learning. So, instead of using a convolutional neural network to approximate action-value function in DQN, a shared convolutional encoder along with a special aggregator&amp;lt;ref&amp;gt;Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas, Dueling Network Architectures for Deep Reinforcement Learning&amp;lt;/ref&amp;gt; is used.&lt;br /&gt;
* Multi-step Learning&lt;br /&gt;
* Distributional RL&lt;br /&gt;
* Noisy Nets&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505623</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505623"/>
		<updated>2018-03-18T06:33:56Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Conclusion */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
[[File:Qlearning table.png|600px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
Although obtaining results with deep reinforcement learning have proven to have low convergence rate, the model proposed in this paper was able to achieve state of the art results on Atari games by using experience replay techniques.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505621</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505621"/>
		<updated>2018-03-18T06:23:10Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Experiments */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
Table 1 shows different scores obtained by different methods on all the games. One of the scores is the one obtained by human players. The proposed model was able to achieve better results than humans in three of the games (Breakout, Enduro, Pong). Other than that, all final scores achieved by the model are higher than the state of the art methods on Atari games. One reason of getting lower scores than humans in some of the games may be the challenging nature of the game: finding a strategy that extends over long time scales.&lt;br /&gt;
[[File:Qlearning table.png|600px||thumb|&#039;&#039;&#039;Table 1.&#039;&#039;&#039; Comparing average total reward for different learning methods. Lower table shows the results obtained by different metrics of the proposed model with last row being the best results.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Qlearning_table.png&amp;diff=505620</id>
		<title>File:Qlearning table.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Qlearning_table.png&amp;diff=505620"/>
		<updated>2018-03-18T06:19:47Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=evaluation table}}&lt;br /&gt;
|date=2018-03-17 23:19:00&lt;br /&gt;
|source=Playing Atari with deep reinforcement learning&lt;br /&gt;
|author=V. Minh, et. al.&lt;br /&gt;
|permission=&lt;br /&gt;
|other_versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{subst:uwl}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Uploaded with UploadWizard]]&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505614</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505614"/>
		<updated>2018-03-18T06:10:48Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Experiments */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
helooooo&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505613</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505613"/>
		<updated>2018-03-18T06:10:25Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Experiments */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505612</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505612"/>
		<updated>2018-03-18T06:09:54Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Experiments */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505611</id>
		<title>Course:CPSC522/Deep Reinforcement Learning</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Deep_Reinforcement_Learning&amp;diff=505611"/>
		<updated>2018-03-18T06:09:10Z</updated>

		<summary type="html">&lt;p&gt;AINAZHAJIMORADLOU: /* Content */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:CPSC522]]&lt;br /&gt;
== Deep Reinforcement Learning ==&lt;br /&gt;
One sentence summary&lt;br /&gt;
&lt;br /&gt;
Principal Author: Ainaz Hajimoradlou&amp;lt;br /&amp;gt;&lt;br /&gt;
Collaborators:&lt;br /&gt;
== Abstract  ==&lt;br /&gt;
This should be a brief summary that tells us what is covered in this page.&lt;br /&gt;
===Builds on===&lt;br /&gt;
Put links to the more general categories that this builds on. These links should all be in sentences that make sense without following the links. You can only rely on technical knowledge that is in these links (and their transitive closure). There is no need to put the transitive closure of the links.&lt;br /&gt;
===Related Pages===&lt;br /&gt;
This should contain the reverse links to the &amp;quot;builds on&amp;quot; links as well as other related pages. Use in sentences.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
=== Playing Atari with Deep Reinforcement Learning ===&lt;br /&gt;
==== Introduction ====&lt;br /&gt;
Throughout past years, deep learning techniques have been widely applied to many vision and speech problems and they were able to improve the final results to large extents. However, these techniques are not very successful when applied to reinforcement learning problems due to variety of reasons:&lt;br /&gt;
* Most of the successful deep learning techniques require a great amount of hand labelled data for training but reinforcement learning algorithms are trained based on a scalar reward which is usually noisy, sparse and delayed. In other words, in order to train an RL algorithm the corresponding reward for each possible action should be received which can take a long time (e.g. thousands of time steps). This is quite challenging compared to the direct relation between targets and inputs in supervised learning.&lt;br /&gt;
* Machine learning and deep learning methods are mostly relied on data being IID. However, most of the data sequences in reinforcement learning are highly correlated.&lt;br /&gt;
* Deep methods are trained based on the assumption that training/testing dataset has an unknown stationary distribution. But in reinforcement learning, distribution of the data changes whenever the model encounters new behaviours in the environment.&lt;br /&gt;
In order to alleviate these problems, experience replay techniques are often used in RL problems. This is typically sampling randomly from the previous experience or transitions. As a result, the training distribution becomes smoother and easier to learn.&lt;br /&gt;
&lt;br /&gt;
This paper proposes a deep reinforcement learning algorithm for 7 different Atari games. No prior knowledge is passed to the agent except the possible actions, reward, terminal signals and the video screen of the play which is available to human players as well.&lt;br /&gt;
&lt;br /&gt;
==== Background ====&lt;br /&gt;
The agent interacts with the environment &amp;lt;math&amp;gt;\varepsilon&amp;lt;/math&amp;gt; in a sequence of actions, observations and rewards. Action &amp;lt;math&amp;gt;a_t&amp;lt;/math&amp;gt; can be chosen from a set of legal actions &amp;lt;math&amp;gt;\Alpha = \{1, ..., K\}&amp;lt;/math&amp;gt;. After doing an action, the internal state of the emulator along with the game score is changed. This change is the reward &amp;lt;math&amp;gt;r_t&amp;lt;/math&amp;gt; at time step &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. The environment can be stochastic or deterministic. As the agent is only able to observe the raw pixels of the screen and not the whole state, the task is partially observable. Therefore, in order to understand the current situation a sequence of actions and observations (&amp;lt;math&amp;gt;x_1, a_1, x_2, a_2, ..., a_{t-1}, x_t&amp;lt;/math&amp;gt;) should be considered. Each sequence is a state (&amp;lt;math&amp;gt;s_t&amp;lt;/math&amp;gt;) at time &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;. States are distinct and have the [https://en.wikipedia.org/wiki/Markov_property Markov property]. Therefore, reinforcement learning techniques can be applied to this problem.&lt;br /&gt;
&lt;br /&gt;
The goal is acting in a way that maximizes future rewards. But Instead of maximizing future rewards, future discounted return (&amp;lt;math&amp;gt;R_t = \sum_{t&#039; = t}^T \gamma^{t&#039; - t} r_{t&#039;}&amp;lt;/math&amp;gt;) is maximized. &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is the finishing time step of the game. One of the reasons of using a discount factor is to avoid the sum of becoming infinite so that it can converge. Optimal action-value function &amp;lt;math&amp;gt;Q^*(s, a)&amp;lt;/math&amp;gt; is defined as maximum expected return by following any strategy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, after seeing sate &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and doing action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \max_{\pi} \mathbb{E}\left[R_t | s_t = s, a_t = a, \pi\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This formula obeys [https://en.wikipedia.org/wiki/Bellman_equation Bellman equation]. Meaning that, if the optimal value &amp;lt;math&amp;gt;Q^*(s&#039;, a&#039;)&amp;lt;/math&amp;gt; of state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt; is known at the next time step for all possible actions &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt;, the optimal strategy is choosing action &amp;lt;math&amp;gt;a&#039;&amp;lt;/math&amp;gt; that maximizes the expected value of &amp;lt;math&amp;gt;r + \gamma Q^*(s&#039;, a&#039;) &amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Q^*(s, a) = \mathbb{E}_{s&#039; \in \varepsilon}\left[r + \gamma \max_{a&#039;} Q^*(s&#039;, a&#039;) | s, a\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, the basic idea of value iteration algorithms is to iteratively update the action-value function by using Bellman equation. This converges to the optimal action-value function as the iterations go to infinity. This is impractical because the function is estimated separately for each state, without generalization. In practice, a function approximator is used to estimate it which can be linear or nonlinear. Neural networks are nonlinear function approximators. If the parameters of the network is &amp;lt;math&amp;gt;\theta&amp;lt;/math&amp;gt;, the loss function at each iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; will be as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.)}\left[(y_i - Q(s, a; \theta_i))^2\right], \quad y_i = \mathbb{E}_{s&#039; \sim \varepsilon}\left[r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1})\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; is the target value for iteration &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The loss function simply tries to minimize the squared error between the predicted value of action-value function given current parameters and target value given previous parameters which is easily computed using bellman equation. &amp;lt;math&amp;gt;\rho(s, a)&amp;lt;/math&amp;gt; is called the behaviour distribution and is simply the distribution over states and actions. The derivative of the loss function is as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\nabla L_i(\theta_i) = \mathbb{E}_{s, a \sim \rho(.) | s&#039; \sim \varepsilon}\left[\left( r + \gamma \max_{a&#039;}Q(s&#039;, a&#039;; \theta_{i-1}) - Q(s, a; \theta_i)\right) \nabla_{\theta_i} Q(s, a; \theta_i)\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In order to make gradient updates less computationally expensive, single samples from the behaviour distribution is used instead of using the full expectation. This distribution is often selected by an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach. Moreover, the algorithm is model free and off policy.&lt;br /&gt;
&lt;br /&gt;
==== Deep Reinforcement Learning ====&lt;br /&gt;
The goal of the paper is doing reinforcement learning using deep networks on Atari games without any preprocessing on the data. It&#039;s worth noting that combining model free reinforcement learning algorithms like Q-learning with non linear approximators or off policy learning can cause the network to diverge. Although this can partially be addressed by gradient temporal difference, the problem of using deep RL still remains a challenging problem. &lt;br /&gt;
&lt;br /&gt;
In order to be able to use machine learning techniques on the problem, an experience replay technique should be used. This is simply storing the previous experiences &amp;lt;math&amp;gt;e_t = (s_t, r_t, s_{t+1})&amp;lt;/math&amp;gt; into a replay memory &amp;lt;math&amp;gt;\mathcal{D}&amp;lt;/math&amp;gt;. The pseudo code of the algorithm is shown below. In each time step of the game, Q-learning updates are applied to random samples of experiences stored in memory. After that, an action is chosen based on an &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;-greedy approach meaning that for probabilities below &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; a random action is chosen instead of the best possible action. &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; is a function that is used to produce fixed length representations for histories.&lt;br /&gt;
&lt;br /&gt;
===== Deep Q-learning with experience replay =====&lt;br /&gt;
 1. Initialize replay memory D to capacity N&lt;br /&gt;
 2. Initialize action-value function Q with random weights &lt;br /&gt;
 3. for episode = 1, M do&lt;br /&gt;
 4.     Initialize sequence s1 = {x1} and preprocessed sequenced φ1 = φ(s1) &lt;br /&gt;
 5.     for t = 1, T do&lt;br /&gt;
 6.         With probability ε select a random action at&lt;br /&gt;
 7.         otherwise select at = maxa Q∗(φ(st), a; θ)&lt;br /&gt;
 8.         Execute action at in emulator and observe reward rt and image xt+1 Set st+1 = st, at, xt+1 and preprocess φt+1 = φ(st+1)&lt;br /&gt;
 9.         Store transition (φt, at, rt, φt+1) in D&lt;br /&gt;
 10.       Sample random minibatch of transitions (φj , aj , rj , φj +1 ) from D&lt;br /&gt;
 11.       if φj+1 is terminal&lt;br /&gt;
 12.           Set yj = rj&lt;br /&gt;
 13.       else&lt;br /&gt;
 14.           Set yj = rj + γ max_a′ Q(φj+1, a′; θ)&lt;br /&gt;
 15.       Perform a gradient descent step on (yj − Q(φj , aj ; θ))^2&lt;br /&gt;
 16.   end for&lt;br /&gt;
 17. end for&lt;br /&gt;
&lt;br /&gt;
This approach has three main advantages compared to approaches such as online Q-learning:&lt;br /&gt;
* It is more efficient as each step of experience is used in many weight updates.&lt;br /&gt;
* The updates are less variant as samples are chosen randomly instead of consecutively and this breaks the correlation between samples to some extent.&lt;br /&gt;
* In on-policy learning, the current parameters of the network determine the next data samples that the parameters should be trained on. As a result, if the maximizing action is to move left then data samples are dominated by the left hand side of the image. This will result in unwanted feedback loops or being stuck in local minima. Experience replay avoids such problems as behaviour distribution is averaged over its previous states, smoothing out learning and avoiding oscillations in parameters.&lt;br /&gt;
The proposed algorithm by the paper uses uniform sampling with a memory buffer of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. However, this approach gives equal importance to all experiences and does not differentiate important transitions. A better sampling technique may be the one that puts more emphasis on transitions from which the model learns the most.&lt;br /&gt;
&lt;br /&gt;
Training on the original pixel images of Atari is computationally expensive. Therefore, a preprocessing stage is used for resizing the images and converting them to grey scale. Besides, there are multiple ways to parametrize Q function using a neural network. This paper uses a convolutional architecture with only the state representation as input and multiple outputs corresponding to each action. Therefore, the network has one forward pass and Q values can be computed for all actions in each state. There are three hidden layers in the network and the final layer is a fully connected layer with a single output for each possible action.&lt;br /&gt;
&lt;br /&gt;
==== Experiments ====&lt;br /&gt;
The model is trained with the same parameters for all different games. The only change is the scaling of the reward function: positive rewards are fixed to 1 and negative rewards are fixed to -1 with 0 being the unchanged rewards. Besides, the paper uses a frame-skipping technique that is instead of choosing actions in each frame, the actions are chosen every k frames and the last action is repeated in skipped frames. This will allow the agent to play roughly k times more without significantly increasing the runtime.&lt;br /&gt;
&lt;br /&gt;
As with any machine learning algorithm, the results should be evaluated. Here, one reasonable choice might be tracking the average total reward metric. But, this tends to be very noisy as the small changes in weights of the network can result in big changes of the distribution. Therefore, a more stable metric can be the estimated value of the action-value function &amp;lt;math&amp;gt;Q&amp;lt;/math&amp;gt; which shows how much discounted reward can be obtained by following a policy in a given state. Figure 1 shows the comparison between these two metrics for evaluation.&lt;br /&gt;
&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;&lt;br /&gt;
[[File:Qlearning eval.png|1200px||thumb|&#039;&#039;&#039;Figure 1.&#039;&#039;&#039; Two figures on the left show the average reward per episode metric while the two on the right show the estimated Q function.&amp;lt;ref name=&amp;quot;paper1&amp;quot;&amp;gt;V. Minh, et al., Playing Atari with Deep Reinforcement Learning&amp;lt;/ref&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
==== Conclusion ====&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
Put your annotated bibliography here. Add links where appropriate.&lt;br /&gt;
{{Reflist|group!=note}}&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1312.5602.pdf V. Minh, et al., Playing Atari with Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
[https://arxiv.org/pdf/1710.02298.pdf M. Hessel, et. al., Rainbow: Combining Improvements in Deep Reinforcement Learning]&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;br /&gt;
&lt;br /&gt;
{{cc-by-nc-sa-3.0}}&lt;/div&gt;</summary>
		<author><name>AINAZHAJIMORADLOU</name></author>
	</entry>
</feed>