<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=TommasoDAmico</id>
	<title>UBC Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.ubc.ca/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=TommasoDAmico"/>
	<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/Special:Contributions/TommasoDAmico"/>
	<updated>2026-08-10T06:00:44Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.9</generator>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=597281</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=597281"/>
		<updated>2020-04-25T19:11:21Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: rearranged future combinations 2020 for marking consistency&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2020|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
== Future combinations (2020) ==&lt;br /&gt;
* M0 [[Course:CPSC522/Grouped Prioritized Experience Replay (GPER)|Grouped Prioritized Experience Replay (GPER)]] &lt;br /&gt;
* M1 [[Course:CPSC522/Alternative_Classifiers|Alternative Classifiers]]&lt;br /&gt;
* [[Course:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations|M2 Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations]]&lt;br /&gt;
* M3 [[Course:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps|Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps]]&lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=597199</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=597199"/>
		<updated>2020-04-24T19:52:57Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about. This extension is motivated by training excersices seen in a variety of sports, where an athlete repeatedly preforms actions in a set of similar envrionment states or performs similar actions in a set of random environment states. The former can be seen in soccer, where the grouped training excersices include training freekicks, cornerballs, counterattacks and other more complicated states that the game presents. The latter can be seen in table tennis, where the grouped training excersices include training a players forehand return( hitting the ball with the red side of the paddle) and backhand return(hitting the ball with the black side of the paddle). This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay.  &lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Temporal Difference Learning &amp;amp; TD Error ====&lt;br /&gt;
Temporal Difference Learning can be more easily understood in comparison to its conterparts  methods such as monte carlo methods. In Temporal Difference Learning, the predictions of current states based on the environment are updated with each new time step, unlike Monte Carlo Methods where these are only updated at the end of an episode. This means that while an agent is actively exploring its environment, the feedback it receives influences its predictions of the current state at each time step. The TD error function reports back the difference between the estimated reward at any given state or time step and the actual reward received. Temporal difference is the basis of algorithms such as Q-learning, TD(0) and SARSA.&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is stored according to its groupping. The TD Error of all stored experiences are then calculated and each group average TD error is obtained. The group with the highest average TD error is then sampled from at uniform.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those experiences will be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
[[File:GPER_Pseudo-code.png|center|thumb|463x463px|Simple implementation of the GPER sampling algorithm]]  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
Wrongful dropoff and pickup have been ket at a low penalty to avoid the agent from learning to avoid the action faster than it would find the reward. This was done as during fine tuning it was found that the agents would often be stuck in local minima where they would just move back around until the end of the episode. &lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (4 groups, each group makes a quadrant of the grid)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The table below summarizes some values from the results:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+Scores of each algorithm over 40000 episodes of Modified Taxi-v3 &lt;br /&gt;
!&lt;br /&gt;
!Q-Learning&lt;br /&gt;
!Experience Replay&lt;br /&gt;
!Proportional PER&lt;br /&gt;
!Rank-Based PER&lt;br /&gt;
!Action GPER&lt;br /&gt;
!Passenger Location GPER&lt;br /&gt;
!Destination Location GPER&lt;br /&gt;
!XY Taxi Location GPER&lt;br /&gt;
|-&lt;br /&gt;
|Average Reward&lt;br /&gt;
|475.49&lt;br /&gt;
|499.17&lt;br /&gt;
|480.06&lt;br /&gt;
|483.38&lt;br /&gt;
| -578.94&lt;br /&gt;
| -539.39&lt;br /&gt;
|183.27&lt;br /&gt;
|355.46&lt;br /&gt;
|-&lt;br /&gt;
|Highest Reward&lt;br /&gt;
|991&lt;br /&gt;
|991&lt;br /&gt;
|990&lt;br /&gt;
|990&lt;br /&gt;
|991&lt;br /&gt;
|991&lt;br /&gt;
|991&lt;br /&gt;
|991&lt;br /&gt;
|-&lt;br /&gt;
|Average Episode Length&lt;br /&gt;
|318.6&lt;br /&gt;
|302.96&lt;br /&gt;
|316.74&lt;br /&gt;
|313.89&lt;br /&gt;
|755.99&lt;br /&gt;
|741.19&lt;br /&gt;
|435.05&lt;br /&gt;
|364.41&lt;br /&gt;
|-&lt;br /&gt;
|Lowest Episode Length&lt;br /&gt;
|9&lt;br /&gt;
|9&lt;br /&gt;
|10&lt;br /&gt;
|10&lt;br /&gt;
|9&lt;br /&gt;
|9&lt;br /&gt;
|9&lt;br /&gt;
|9&lt;br /&gt;
|} &lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested. &lt;br /&gt;
&lt;br /&gt;
Given the rewards of -1 for movement and wrongful actions and +1000 for completing the task the minimum score is of -900 and the maximum will approach +1000. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for Rewards over time.png|center|thumb|1197x1197px|Average rewards over 40000 episodes for each tested algorithm]] &lt;br /&gt;
&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time. In the plot below lower scores are better as it represents the average number of movements it takes for an agent to complete the game over the learning process. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for Episode Length over time.png|center|thumb|1189x1189px|Average Epsiode Length over 40000 episodes for each tested algorithm]]&lt;br /&gt;
After plotting these, there seemed to be some ambiguity on the variance of these averages, do curves such as Action GPER actually understand the game but only complete it 1/4 of the time or are they taking 750 steps every episode to complete the tasks? Below is a curve looking at the change in episode lenght over time. This has been further smoothed out to facilitate visualization of the trends.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for change in rewards over time.png|center|thumb|1188x1188px|Average Change in Rewards over 40000 episodes for each algorithm tested]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
From the first plot shown it can be seen that although all algorithms seem to converge, some converge at lower scores than others. As emphasized in the thrid plot this is mainly due to the averaging effect used to clean up the data for visualization. Meaning that most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that most episodes they will not successfully terminate the game and are more likely to exert erratic behaviours. &lt;br /&gt;
&lt;br /&gt;
It is possible that with further training these would eventually converge to the global maximum although it appears that they have settled in local maxima. Tuning of the epsylon greedy parameter may help in this case as it would allow these algorithms to explore further scenarious that they seem to not have been exposed to prior. &lt;br /&gt;
&lt;br /&gt;
It can be seen that of the GPER methods tested achieved lower performance compared to prior established algorithms. Provided that the same parameters were used for all trainings and there are no extra parameters introduced in the proposed adaptation, the conclusion may be that the technique does not compare in performance to prior established algorithms. &lt;br /&gt;
&lt;br /&gt;
From the second plot illustrating episode length over time, the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger. &lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, where the grid is divided into 4 quadrants/groups. Unlike passenger location and destination location grouping, XY grouping is useful throughout the entire game, which could have lead to its increased performance.However, action groupings, which are useful throughout the game as well, did not perform as well. This leads to the idea that state grouping could generally be better than action grouping, nonetheless, further testing on different environment and parameter fine tuning should be performed before coming to such a conclusion. &lt;br /&gt;
&lt;br /&gt;
It is interesting to notice that initially XY Taxi Location GPER also seems to perform on par with the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
In the third plot we see the change in rewards of each algorithm over their learning process. IT can be seen that established algoriths initially have a high variance in their rewards, but converge to have small differences, indicating they have perfected their strategy. The same cannot be said about the GPER algorithms where large differences in score remain constant until the end of the learning process. This indicates that these did not converge to the solution and still only manage to complete the task occasionally.&lt;br /&gt;
&lt;br /&gt;
=== Conclusion and Further Work ===&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above there is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
Further testing of these algorithms is required with different environments and paramenters to determine their functionality and usefulness. As shown in the results above, little difference among pre established algorithms can be seen in this particular example, demonstrating that perhaps more eleborate environments are required to highlight their differences.&lt;br /&gt;
&lt;br /&gt;
An aspect that can be further looked at is that of utilizing probability among groups when sampling instead of sampling from the grouping with the worst error every time, this may avoid algorithms getting stuck on a specific issue that may be learned by learning other aspects of the overall task.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=597192</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=597192"/>
		<updated>2020-04-24T19:06:33Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about. This extension is motivated by training excersices seen in a variety of sports, where an athlete repeatedly preforms actions in a set of similar envrionment states or performs similar actions in a set of random environment states.The former can be seen in soccer, where the grouped training excersices include training freekicks, cornerballs, counterattacks and other more complicated states that the game presents. The latter can be seen in table tennis, where the grouped training excersices include training a players forehand return( hitting the ball with the red side of the paddle) and backhand return(hitting the ball with the black side of the paddle). This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay.  &lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is stored according to its groupping. The TD Error of all stored experiences are then calculated and each group average TD error is obtained. The group with the highest average TD error is then sampled from at uniform.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those experiences will be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
[[File:GPER_Pseudo-code.png|center|thumb|463x463px|Simple implementation of the GPER sampling algorithm]]  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
Wrongful dropoff and pickup have been ket at a low penalty to avoid the agent from learning to avoid the action faster than it would find the reward. This was done as during fine tuning it was found that the agents would often be stuck in local minima where they would just move back around until the end of the episode. &lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (4 groups, each group makes a quadrant of the grid)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested.  &lt;br /&gt;
&lt;br /&gt;
Given the rewards of -1 for movement and wrongful actions and +1000 for completing the task the minimum score is of -900 and the maximum will approach +1000. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for Rewards over time.png|center|thumb|1180.99x1180.99px|Average rewards over 40000 episodes for each tested algorithm]] &lt;br /&gt;
&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time. In the plot below lower scores are better as it represents the average number of movements it takes for an agent to complete the game over the learning process. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for Episode Length over time.png|center|thumb|1188.99x1188.99px|Average Epsiode Length over 40000 episodes for each tested algorithm]]&lt;br /&gt;
After plotting these, there seemed to be some ambiguity on the variance of these averages, do curves such as Action GPER actually understand the game but only complete it 1/4 of the time or are they taking 750 steps every episode to complete the tasks? Below is a curve looking at the change in episode lenght over time. This has been further smoothed out to facilitate visualization of the trends.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Testing Results for change in rewards over time.png|center|thumb|1188x1188px|Average Change in Rewards over 40000 episodes for each algorithm tested]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
From the first plot shown it can be seen that although all algorithms seem to converge, some converge at lower scores than others. As emphasized in the thrid plot this is mainly due to the averaging effect used to clean up the data for visualization. Meaning that most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that most episodes they will not successfully terminate the game and are more likely to exert erratic behaviours. &lt;br /&gt;
&lt;br /&gt;
It is possible that with further training these would eventually converge to the global maximum although it appears that they have settled in local maxima. Tuning of the epsylon greedy parameter may help in this case as it would allow these algorithms to explore further scenarious that they seem to not have been exposed to prior. &lt;br /&gt;
&lt;br /&gt;
It can be seen that of the GPER methods tested achieved lower performance compared to prior established algorithms. Provided that the same parameters were used for all trainings and there are no extra parameters introduced in the proposed adaptation, the conclusion may be that the technique does not compare in performance to prior established algorithms. &lt;br /&gt;
&lt;br /&gt;
From the second plot illustrating episode length over time, the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger. &lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, where the grid is divided into 4 quadrants/groups. Unlike passenger location and destination location grouping, XY grouping is useful throughout the entire game, which could have lead to its increased performance.However, action groupings, which are useful throughout the game as well, did not perform as well. This leads to the idea that state grouping could generally be better than action grouping, nonetheless, further testing on different environment and parameter fine tuning should be performed before coming to such a conclusion. &lt;br /&gt;
&lt;br /&gt;
It is interesting to notice that initially XY Taxi Location GPER also seems to perform on par with the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
In the third plot we see the change in rewards of each algorithm over their learning process. IT can be seen that established algoriths initially have a high variance in their rewards, but converge to have small differences, indicating they have perfected their strategy. The same cannot be said about the GPER algorithms where large differences in score remain constant until the end of the learning process. This indicates that these did not converge to the solution and still only manage to complete the task occasionally.&lt;br /&gt;
&lt;br /&gt;
=== Conclusion and Further Work ===&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above there is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
Further testing of these algorithms is required with different environments and paramenters to determine their functionality and usefulness. As shown in the results above, little difference among pre established algorithms can be seen in this particular example, demonstrating that perhaps more eleborate environments are required to highlight their differences.&lt;br /&gt;
&lt;br /&gt;
An aspect that can be further looked at is that of utilizing probability among groups when sampling instead of sampling from the grouping with the worst error every time, this may avoid algorithms getting stuck on a specific issue that may be learned by learning other aspects of the overall task.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_change_in_rewards_over_time.png&amp;diff=597190</id>
		<title>File:Grouped Prioritized Experience Replay Testing Results for change in rewards over time.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_change_in_rewards_over_time.png&amp;diff=597190"/>
		<updated>2020-04-24T18:12:28Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Grouped Prioritized Experience Replay Testing Results for change in rewards over time}}&lt;br /&gt;
|date=2020-04-24&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_Episode_Length_over_time.png&amp;diff=597189</id>
		<title>File:Grouped Prioritized Experience Replay Testing Results for Episode Length over time.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_Episode_Length_over_time.png&amp;diff=597189"/>
		<updated>2020-04-24T18:12:28Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Grouped Prioritized Experience Replay Testing Results for Episode Length over time}}&lt;br /&gt;
|date=2020-04-24&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_Rewards_over_time.png&amp;diff=597188</id>
		<title>File:Grouped Prioritized Experience Replay Testing Results for Rewards over time.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Testing_Results_for_Rewards_over_time.png&amp;diff=597188"/>
		<updated>2020-04-24T18:12:28Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Grouped Prioritized Experience Replay Testing Results for Rewards over time}}&lt;br /&gt;
|date=2020-04-24&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Improving_Prediction_Accuracy_of_User_Cognitive_Abilities_for_User-Adaptive_Narrative_Visualizations/Feedback_(2)&amp;diff=597040</id>
		<title>Thread:Course talk:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations/Feedback (2)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Improving_Prediction_Accuracy_of_User_Cognitive_Abilities_for_User-Adaptive_Narrative_Visualizations/Feedback_(2)&amp;diff=597040"/>
		<updated>2020-04-23T15:42:43Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: New thread: Feedback&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Great page, very readable! Same as the critiques below, I cna&#039;t find the links to the study you reference. Perhaps consider retitling your component of the work from extension, I understood the meaning after looking back at the wiki structure but was confused at first. &lt;br /&gt;
Do you mean Logistic regression for logic regression? I wasn&#039;t able to double check this.&lt;br /&gt;
Overall really nice page&lt;br /&gt;
&lt;br /&gt;
The topic is relevant for the course. 5 The writing is clear and the English is good. 5 The page is written at an appropriate level for CPSC 522 students. 5 The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 5 The abstract is a concise and clear summary. 5 There were appropriate (original) examples that helped make the topic clear.- There was appropriate use of (pseudo-) code. - It had a good coverage of representations, semantics, inference and learning. 5 It is correct. 5 It was neither too short nor too long for the topic. 5 It was an appropriate unit for a page. 5 It links to appropriate other pages in the wiki. 1 The references and links to external pages are well chosen. - I would recommend this page to someone who wanted to find out about the topic. 5 This page should be highlighted as an exemplary page for others to emulate. 4 If I was grading it out of 20, I would give it: 18&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Exploring_Results_of_Conditional_Generative_Adversarial_Networks_with_Self-Organizing_Maps/Feedback_(3)&amp;diff=597039</id>
		<title>Thread:Course talk:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps/Feedback (3)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Exploring_Results_of_Conditional_Generative_Adversarial_Networks_with_Self-Organizing_Maps/Feedback_(3)&amp;diff=597039"/>
		<updated>2020-04-23T15:32:11Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: New thread: Feedback&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Interesting topic and well written, I found it a bit confusing to read the methodology after the experiment, I think reorganizing it might help your structuring. I also strugled a bit interpreting the heat maps, perhaps plaing them next to the image used to formulate it might make it more interpretable. Overall pretty good!&lt;br /&gt;
&lt;br /&gt;
The topic is relevant for the course. 5 The writing is clear and the English is good. 5 The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 5 The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4 The abstract is a concise and clear summary. 4 There were appropriate (original) examples that helped make the topic clear. 4 There was appropriate use of (pseudo-) code. It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). 5 It is correct. 5 It was neither too short nor too long for the topic. 5 It was an appropriate unit for a page. 5 It links to appropriate other pages in the wiki. 5 The references and links to external pages are well chosen. 5 I would recommend this page to someone who wanted to find out about the topic. 5 This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
&lt;br /&gt;
If I was grading it out of 20, I would give it: 18&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Alternative_Classifiers/Feedback_(2)&amp;diff=597038</id>
		<title>Thread:Course talk:CPSC522/Alternative Classifiers/Feedback (2)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Alternative_Classifiers/Feedback_(2)&amp;diff=597038"/>
		<updated>2020-04-23T15:18:18Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: New thread: Feedback&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The topic is interesting and movel to explore from my understanding. Although I believe the use of language may be slightly advanced for this course, further explaining certain core concepts might help the readability of the page. Otherwise I think the page looks good, great choice.&lt;br /&gt;
&lt;br /&gt;
The topic is relevant for the course. 5 The writing is clear and the English is good. 5 The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds). 4 The formalism (definitions, mathematics) was well chosen to make the page easier to understand. 4 The abstract is a concise and clear summary. 4 There were appropriate (original) examples that helped make the topic clear. 4 There was appropriate use of (pseudo-) code. It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic). 5 It is correct. 5 It was neither too short nor too long for the topic. 5 It was an appropriate unit for a page. 5 It links to appropriate other pages in the wiki.  The references and links to external pages are well chosen. 5 I would recommend this page to someone who wanted to find out about the topic. 5 This page should be highlighted as an exemplary page for others to emulate. 4&lt;br /&gt;
&lt;br /&gt;
If I was grading it out of 20, I would give it: 18&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596874</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596874"/>
		<updated>2020-04-21T21:27:32Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2020|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
== Future combinations (2020) ==&lt;br /&gt;
* [[Course:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations|Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations]]&lt;br /&gt;
* M3 [[Course:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps|Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps]]&lt;br /&gt;
* M1 [[Course:CPSC522/Alternative_Classifiers|Alternative Classifiers]] (2020)&lt;br /&gt;
* M0 [[Course:CPSC522/Grouped Prioritized Experience Replay (GPER)|Grouped Prioritized Experience Replay (GPER)]] (2020)&lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596839</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596839"/>
		<updated>2020-04-21T14:52:37Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about. This extension is motivated by training excersices seen in a variety of sports, where an athlete repeatedly preforms actions in a set of similar envrionment states or performs similar actions in a set of random environment states.The former can be seen in soccer, where the grouped training excersices include training freekicks, cornerballs, counterattacks and other more complicated states that the game presents. The latter can be seen in table tennis, where the grouped training excersices include training a players forehand return( hitting the ball with the red side of the paddle) and backhand return(hitting the ball with the black side of the paddle). This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay.  &lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is stored according to its groupping. The TD Error of all stored experiences are then calculated and each group average TD error is obtained. The group with the highest average TD error is then sampled from at uniform.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those experiences will be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
[[File:GPER_Pseudo-code.png|center|thumb|463x463px|Simple implementation of the GPER sampling algorithm]]  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (4 groups, each group makes a quadrant of the grid)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1183x1183px|Average Rewards over 40000 episodes for each algorithm]]&lt;br /&gt;
The following plot is a zoomed in version of the prior to demonstrate more closely the performance of each algorithm when converging.&lt;br /&gt;
[[File:Zoomed in Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1180x1180px|Zoomed in view of reward plot]]&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Episode Length Plot.png|center|thumb|1185x1185px|Average episode legth over time (max episode length: 900)]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
In the average reward plot it can be seen that although all algorithms seem to converge, some converge at lower scores than others. This is mainly due to the averaging effect used to clean up the data for visualization. most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that the variance in the performance appeared much larger. It is possible that with further training these would eventually converge to the global maximum although it appears that some have settled in local maxima, such as Experience Replay.&lt;br /&gt;
&lt;br /&gt;
It can be seen also that none of the GPER methods tested achieved the same performance as prior established algorithms.&lt;br /&gt;
&lt;br /&gt;
From the last plot illustrating episode length over time,the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger.&lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, where the grid is divided into 4 quadrants/groups. Unlike passenger location and destination location grouping, XY grouping is useful throughout the entire game, which could have lead to its increased performance. However, action groupings, which are useful throughout the game as well, did not perform as well. This leads to the idea that state grouping could generally be better than action grouping, nonetheless, further testing on different environment and parameter fine tuning should be performed before coming to such a conclusion. It is interesting to notice that initially XY Taxi Location GPER also seems to outperform the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above there is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596785</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596785"/>
		<updated>2020-04-20T20:23:04Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (400 groups)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1183x1183px|Average Rewards over 40000 episodes for each algorithm]]&lt;br /&gt;
The following plot is a zoomed in version of the prior to demonstrate more closely the performance of each algorithm when converging.&lt;br /&gt;
[[File:Zoomed in Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1180x1180px|Zoomed in view of reward plot]]&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Episode Length Plot.png|center|thumb|1185x1185px|Average episode legth over time (max episode length: 900)]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
In the average reward plot it can be seen that although all algorithms seem to converge some converge at lower scores than others. This is mainly due to the averaging effect used to clean up the data for visualization. most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that the variance in the performance appeared much larger. It is possible that with further training these would eventually converge to the global maximum although it appears that some have settled in local maxima, such as Experience Replay.&lt;br /&gt;
&lt;br /&gt;
It can be seen also that none of the GPER methods tested achieved the same performance as prior established algorithms.&lt;br /&gt;
&lt;br /&gt;
From the last plot illustrating episode length over time,the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger.&lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, which is the grouping algorithm with the most groups of the ones tested (400 groups). This leads to the idea that reducing the number of elements to prioritize amongst leads to worse performance and that ultimately grouping may not be a valuable addition to the algorithm. This can be argued although further testing and parameter fine tuning should be performed before coming to such a conclusion. It is interesting to notice that initially XY Taxi Location GPER also seems to outperform the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above ther is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596784</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596784"/>
		<updated>2020-04-20T20:22:42Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (400 groups)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1183x1183px|Average Rewards over 40000 episodes for each algorithm]]&lt;br /&gt;
The following plot is a zoomed in version of the prior to demonstrate more closely the performance of each algorithm when converging.&lt;br /&gt;
[[File:Zoomed in Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1180x1180px|Zoomed in view of reward plot]]&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Episode Length Plot.png|center|thumb|1180.99x1180.99px|Average episode legth over time (max episode length: 900)]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
In the average reward plot it can be seen that although all algorithms seem to converge some converge at lower scores than others. This is mainly due to the averaging effect used to clean up the data for visualization. most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that the variance in the performance appeared much larger. It is possible that with further training these would eventually converge to the global maximum although it appears that some have settled in local maxima, such as Experience Replay.&lt;br /&gt;
&lt;br /&gt;
It can be seen also that none of the GPER methods tested achieved the same performance as prior established algorithms.&lt;br /&gt;
&lt;br /&gt;
From the last plot illustrating episode length over time,the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger.&lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, which is the grouping algorithm with the most groups of the ones tested (400 groups). This leads to the idea that reducing the number of elements to prioritize amongst leads to worse performance and that ultimately grouping may not be a valuable addition to the algorithm. This can be argued although further testing and parameter fine tuning should be performed before coming to such a conclusion. It is interesting to notice that initially XY Taxi Location GPER also seems to outperform the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above ther is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596783</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596783"/>
		<updated>2020-04-20T20:21:58Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping (6 groups)&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping (400 groups)&lt;br /&gt;
* Q-Learning with GPER destination location grouping (8 groups)&lt;br /&gt;
* Q-Learning with GPER passenger location grouping (9 groups)&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
For the results below the algorithms have played through 40000 games of Modified Taxi-v3 with a limit of 900 turns per game.&lt;br /&gt;
&lt;br /&gt;
The following plot represent the reward obtained by the agent at each episode averaged every 100 episodes for each algorithm tested. &lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1478.98x1478.98px|Average Rewards over 40000 episodes for each algorithm]]&lt;br /&gt;
The following plot is a zoomed in version of the prior to demonstrate more closely the performance of each algorithm when converging.&lt;br /&gt;
[[File:Zoomed in Grouped Prioritized Experience Replay Reward Plot.png|center|thumb|1474.96x1474.96px|Zoomed in view of reward plot]]&lt;br /&gt;
The following plot demonstrates the average episode length for each algorithm over time.&lt;br /&gt;
[[File:Grouped Prioritized Experience Replay Episode Length Plot.png|center|thumb|1473.98x1473.98px|Average episode legth over time (max episode length: 900)]]&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
In the average reward plot it can be seen that although all algorithms seem to converge some converge at lower scores than others. This is mainly due to the averaging effect used to clean up the data for visualization. most algorithms seem to have reached the maximum highscores although those illustrated with lower averages tended to be more erratic in behaviour, meaning that the variance in the performance appeared much larger. It is possible that with further training these would eventually converge to the global maximum although it appears that some have settled in local maxima, such as Experience Replay.&lt;br /&gt;
&lt;br /&gt;
It can be seen also that none of the GPER methods tested achieved the same performance as prior established algorithms.&lt;br /&gt;
&lt;br /&gt;
From the last plot illustrating episode length over time,the tested GPER algorithms never quite learned the game, appearing to not complete most times. For example in Passenger Location GPER the agent seems to run out of steps about 3 out of 4 times it attempts the game. This may be as the passenger location grouping is only relevant to the game for the first component where the agent needs to find the passenger.&lt;br /&gt;
&lt;br /&gt;
The best performing GPER grouping tested seemed to be grouping by XY Taxi coordinates, which is the grouping algorithm with the most groups of the ones tested (400 groups). This leads to the idea that reducing the number of elements to prioritize amongst leads to worse performance and that ultimately grouping may not be a valuable addition to the algorithm. This can be argued although further testing and parameter fine tuning should be performed before coming to such a conclusion. It is interesting to notice that initially XY Taxi Location GPER also seems to outperform the established algorithms. This may be because of the desired effect of grouping seeked by this experiment, where by grouping experiences by relevant states, the algorithm restricts prioritization only to groups considered relevant by the designer, reducing the need for the algorithm to train against redundant behaviours.&lt;br /&gt;
&lt;br /&gt;
Overall GPER does not seem to be a better algorithm than currently used algorithms for experience replay in Q-Learning. It does although seems to provide valuable insights to which state types may be relevant for the learning task. From the graphs above ther is a clear difference in performance for each grouping which suggests a difference in relevance of each state type to the learning process. This may be relevant to understanding what influences the learning process of an agent over time.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Episode_Length_Plot.png&amp;diff=596774</id>
		<title>File:Grouped Prioritized Experience Replay Episode Length Plot.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Episode_Length_Plot.png&amp;diff=596774"/>
		<updated>2020-04-20T18:53:39Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Plot comparing Grouped Prioritized Experience Replay with other Learning Algorithms based on episode training length over 40000 episodes of modified version of Taxi-v3 OpenAi Gym Game Environment.&lt;br /&gt;
Maximum training length was 900 episodes.}}&lt;br /&gt;
|date=2020-04-20&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Zoomed_in_Grouped_Prioritized_Experience_Replay_Reward_Plot.png&amp;diff=596773</id>
		<title>File:Zoomed in Grouped Prioritized Experience Replay Reward Plot.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Zoomed_in_Grouped_Prioritized_Experience_Replay_Reward_Plot.png&amp;diff=596773"/>
		<updated>2020-04-20T18:53:39Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Zoome in version of plot comparing Grouped Prioritized Experience Replay with other Learning Algorithms based on rewards over 40000 episodes of modified version of Taxi-v3 OpenAi Gym Game Environment.}}&lt;br /&gt;
|date=2020-04-20&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Reward_Plot.png&amp;diff=596772</id>
		<title>File:Grouped Prioritized Experience Replay Reward Plot.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Grouped_Prioritized_Experience_Replay_Reward_Plot.png&amp;diff=596772"/>
		<updated>2020-04-20T18:53:39Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Plot comparing Grouped Prioritized Experience Replay with other Learning Algorithms based on rewards over 40000 episodes of modified version of Taxi-v3 OpenAi Gym Game Environment.}}&lt;br /&gt;
|date=2020-04-20&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596687</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596687"/>
		<updated>2020-04-19T23:24:35Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
&lt;br /&gt;
=== Discussion ===&lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596686</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596686"/>
		<updated>2020-04-19T23:23:55Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Builds on&#039;&#039;&#039; ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Related Pages&#039;&#039;&#039; ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== Background ===&lt;br /&gt;
&lt;br /&gt;
==== &#039;&#039;&#039;Experience Replay&#039;&#039;&#039; ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== &#039;&#039;&#039;Prioritized Experience Replay&#039;&#039;&#039; ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== Hypothesis ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== Testing ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== &#039;&#039;&#039;Results&#039;&#039;&#039; ====&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Discussion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Conclusion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596685</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596685"/>
		<updated>2020-04-19T23:17:02Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Introduction&#039;&#039;&#039; ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Background&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Hypothesis&#039;&#039;&#039; ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== Algorithm ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a type of prioritization focused on groups rather than individual experiences. In order to acomplish group prioritization, after each time step, the memory of the last experience is set according to its groupping. The TD Error of all stored experiences is then calculated and the group average is obtained. The prioritization of the experience replay is therefore executed based on the scoring of the average TD Error of the groupping.&lt;br /&gt;
&lt;br /&gt;
For example, if grouping is done by action, after a certain time step, if the agent moved up, the resulting experience will be placed under the action group &amp;quot;move up&amp;quot;. The  error of each experience is then calculated and averaged for each group. It may be that the agent is most uncertain of movements to the left, therefore those will be more likely to be replayed, ideally improving its performance &amp;quot;going left&amp;quot; more quickly.  &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Testing&#039;&#039;&#039; ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Discussion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Conclusion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596683</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596683"/>
		<updated>2020-04-19T22:37:41Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Introduction&#039;&#039;&#039; ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Background&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Hypothesis&#039;&#039;&#039; ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Testing&#039;&#039;&#039; ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -1 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Discussion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Conclusion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596682</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596682"/>
		<updated>2020-04-19T22:36:55Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2020|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
* Term Project: [[Course:CPSC522/Alternative_Classifiers|Alternative Classifiers]] (2020)&lt;br /&gt;
* M0 [[Course:CPSC522/Grouped Prioritized Experience Replay (GPER)|Grouped Prioritized Experience Replay (GPER)]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
== Future combinations (2020) ==&lt;br /&gt;
* [[Course:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations|Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations]]&lt;br /&gt;
* [[Course:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps|Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps]]&lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596681</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596681"/>
		<updated>2020-04-19T22:23:16Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2020|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
* Term Project: [[Course:CPSC522/Alternative_Classifiers|Alternative Classifiers]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
* M0 [[Course:CPSC522/Grouped Prioritized Experience Replay (GPER)]] (2020)&lt;br /&gt;
&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
== Future combinations (2020) ==&lt;br /&gt;
* [[Course:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations|Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations]]&lt;br /&gt;
* [[Course:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps|Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps]]&lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596680</id>
		<title>Course:CPSC522/Grouped Prioritized Experience Replay (GPER)</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Grouped_Prioritized_Experience_Replay_(GPER)&amp;diff=596680"/>
		<updated>2020-04-19T22:16:13Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: Created page with &amp;quot;== Grouped Prioritized Experience Replay (GPER) == This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Introduction&#039;&#039;&#039; ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Background&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Hypothesis&#039;&#039;&#039; ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Testing&#039;&#039;&#039; ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -10 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful pickup&lt;br /&gt;
* +1000 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Discussion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Conclusion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596679</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=596679"/>
		<updated>2020-04-19T22:15:33Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2020|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
* Term Project: [[Course:CPSC522/Alternative_Classifiers|Alternative Classifiers]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Grouped Prioritized Experience Replay (GPER)]] (2020)&lt;br /&gt;
&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
== Future combinations (2020) ==&lt;br /&gt;
* [[Course:CPSC522/Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations|Improving Prediction Accuracy of User Cognitive Abilities for User-Adaptive Narrative Visualizations]]&lt;br /&gt;
* [[Course:CPSC522/Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps|Exploring Results of Conditional Generative Adversarial Networks with Self-Organizing Maps]]&lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Sandbox:Grouped_Prioritized_Experience_Replay&amp;diff=596675</id>
		<title>Sandbox:Grouped Prioritized Experience Replay</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Sandbox:Grouped_Prioritized_Experience_Replay&amp;diff=596675"/>
		<updated>2020-04-19T22:04:33Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: Created page with &amp;quot;== Grouped Prioritized Experience Replay (GPER) == This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Grouped Prioritized Experience Replay (GPER) ==&lt;br /&gt;
This page explores the concept of grouping prioritized experience replay memories to improve learning efficiency in temporal difference reinforcement learning environments.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Obada Alhumsi, Tommaso D’Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state action pairs. The page below explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
Grouped Prioritized Experience Replay builds upon [https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Prioritized Experience Replay] which in turn builds upon the concept of Experience Replay in [https://wiki.ubc.ca/Course:CPSC522/Deep_Q-Learning Temporal Difference Learning]. The algorithms will use [[Course:CPSC522/Reinforcement Learning with Function approximation|Q-Learning]] as a fundamental building block.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
[https://wiki.ubc.ca/Course:CPSC522/Deep_Q_Network Deep-Q Networks] are used instead of Q-Learning to apply these algorithms in continuos learning environments.&lt;br /&gt;
&lt;br /&gt;
Experience Replay is also tightly related to [https://en.wikipedia.org/wiki/Hippocampal_replay Hippocampal Replay] in the literature.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Introduction&#039;&#039;&#039; ===&lt;br /&gt;
Grouped Prioritized Experience Replay is a hypothetical extension of Prioritized Experience Replay where prioritization is done by grouped experiences rather than individual state-action pairs. This project explores the effectiveness of this method with respect to Q-Learning, Q-Learning with experience replay and Q-Learning with prioritized experience replay. &lt;br /&gt;
&lt;br /&gt;
The idea is to build upon prioritized experience replay by prioritizing groups of experiences rather than individual experiences. The intention is to render learning more efficient by prioritizing experience replay to specific action or state groups with which the agent is most uncertain about.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Background&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
==== Experience Replay ====&lt;br /&gt;
Experience Replay is an extension of Q-learning where instead of just running Q-learning on state-action pairs as they occur, the agent stores these state-actions pair experiences and updates them based on future experiences that occur, essentially introducing the concept of memory to the algorithm. This way the algorithm has more efficient use of its experiences and is able to learn faster than standard Q-learning. &lt;br /&gt;
&lt;br /&gt;
==== Prioritized Experience Replay ====&lt;br /&gt;
Prioritized Experience Replay attempts to optimize Experience Replay to areas where the agent is more uncertain of its actions by prioritizing experiences with a high temporal-difference error (with some stochastic elements). This allows the agent to learn more quickly by avoiding going over experiences it has mastered and focusing on the ones where it seems to be making more mistakes. By doing so Prioritized experience replay allows for more valuable memory storage as only relevant experiences remain stored and are replayed. For more detailed information on this topic visit the page on Deep-Q Networks. &lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Hypothesis&#039;&#039;&#039; ===&lt;br /&gt;
The hypothesis to be tested is that the concept of grouping priorities in experience replay by state or action can improve training efficiency and minimize error for temporal difference learning algorithms in comparison to plain Q-Learning, Q-Learning with Experience Replay and Q-Learning with Prioritized Experience Replay.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Testing&#039;&#039;&#039; ===&lt;br /&gt;
[[File:Taxi-v3.png|thumb|Taxi-v3 Game Environment Render]]&lt;br /&gt;
&lt;br /&gt;
==== Test Environment ====&lt;br /&gt;
The testing for this algorithm was done using a modified version of the [https://gym.openai.com/envs/Taxi-v3/ Taxi-v3 game] environment present in the [https://gym.openai.com/envs/#toy_text OpenAi Gym Library].&lt;br /&gt;
[[File:Taxi-v3 full.png|thumb|Taxi-v3 Game with Passenger inside Taxi]]&lt;br /&gt;
In the original Taxi-v3 game and in the modified version the taxi starts off at a random square and the passenger is at a random location. The taxi drives to the passenger&#039;s location, picks up the passenger, drives to the passenger&#039;s destination and then drops off the passenger. Once the passenger is dropped off, the episode ends.&lt;br /&gt;
&lt;br /&gt;
The following are the labels demonstrated in the render of the game:&lt;br /&gt;
*    letters and numbers (R, G, B, Y, 0-7): locations for passenger and destinations&lt;br /&gt;
*    blue: passenger&lt;br /&gt;
*    magenta: destination&lt;br /&gt;
*    yellow: empty taxi&lt;br /&gt;
*    green: full taxi&lt;br /&gt;
[[File:Taxi mod big.png|thumb|Modified Taxi-v3 game environment]]&lt;br /&gt;
The modified version of the game is an extention of the map and an increase in locations and destinations for the passanger. This was done in an effort to obtain a larger discrete environment to train the algorithms on to more suitably display the differences in training efficiency.&lt;br /&gt;
&lt;br /&gt;
The original Taxi-v3 game has 500 states (5 x-axis positions, 5 y-axis positions, 4 destination locations and 5 passanger locations) and 6 actions (move up, move down, move left, move right, pickup passenger and dropoff passenger).&lt;br /&gt;
&lt;br /&gt;
The Modified Taxi game has 28800 states (20 x-axis positions, 20 y-axis positions, 8 destination locations and 9 passanger locations) and the same 6 actions as the original.&lt;br /&gt;
&lt;br /&gt;
Rewards in the modified Taxi game were given as follows:&lt;br /&gt;
* -1 for movement&lt;br /&gt;
* -10 for wrong dropoff or pickup&lt;br /&gt;
* +100 for successful pickup&lt;br /&gt;
* +1000 for successful dropoff&lt;br /&gt;
&lt;br /&gt;
==== Methodology ====&lt;br /&gt;
In order to test the performance of Grouped Prioritized Experience Replay a set of algorithms was used as baseline for performance. Since GPER builds upon other concepts the following algorithms were used in comparison:&lt;br /&gt;
* Q-Learning&lt;br /&gt;
* Q-Learning with Experience Replay&lt;br /&gt;
* Q-Learning with Proportional Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with Rank-Based Prioritized Experience Replay&lt;br /&gt;
* Q-Learning with GPER action grouping&lt;br /&gt;
* Q-Learning with GPER x-y position state grouping&lt;br /&gt;
* Q-Learning with GPER destination location grouping&lt;br /&gt;
* Q-Learning with GPER passenger location grouping&lt;br /&gt;
All algorithms were trained with the same parameters and in the environment to avoid biases.&lt;br /&gt;
&lt;br /&gt;
==== Results ====&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Discussion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Conclusion&#039;&#039;&#039; ===&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Taxi-v3_full.png&amp;diff=596672</id>
		<title>File:Taxi-v3 full.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Taxi-v3_full.png&amp;diff=596672"/>
		<updated>2020-04-19T21:45:18Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Taxi-v3 OpenAi Gym game Illustration with passenger on taxi}}&lt;br /&gt;
|date=2020-04-19&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Taxi-v3.png&amp;diff=596671</id>
		<title>File:Taxi-v3.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Taxi-v3.png&amp;diff=596671"/>
		<updated>2020-04-19T21:45:18Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Taxi-v3 OpenAi Gym game Illustration}}&lt;br /&gt;
|date=2020-04-19&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:Taxi_mod_big.png&amp;diff=596670</id>
		<title>File:Taxi mod big.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:Taxi_mod_big.png&amp;diff=596670"/>
		<updated>2020-04-19T21:45:18Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Modified Taxi-v3 OpenAi Gym game}}&lt;br /&gt;
|date=2020-04-19&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586706</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586706"/>
		<updated>2020-03-14T18:43:13Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)|623x623px]]&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment&amp;lt;ref&amp;gt;Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
*&amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
After executing one time step an agent needs to update its belief of the state of the environment it is in&amp;lt;ref name=&amp;quot;:0&amp;quot;&amp;gt;Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. The belief update is an actionless step and is therefore the same as seen on a [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Model] belief update. &lt;br /&gt;
&lt;br /&gt;
The belief state defined as &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; is a function over all states &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; from [0,1] that sums up to 1. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
[[File:DroneSurveilance.gif|thumb|Drone Navigating around space avoiding ground agent using POMDP|439x439px]]&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
An example of a POMDP in action can be seen on the left, this animation illustrates a drone (in yellow) navigating a tiled environment. Its objective is to reach the oposing green square without coming into contact with the ground agent (in red). The blue squares represent the drone&#039;s observable environment. Since the drone is not capale of observing the state of the ground agent at all times, MDP is not tractable for this task, POMDP enables it to estimate a probability distribution of the likely states the ground agent might be in and act accordingly.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the original POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
[[File:POMDP policy.png|thumb|POMDP decomposed into a state estimator and a policy.]]&lt;br /&gt;
&lt;br /&gt;
=== Policy and Value Function ===&lt;br /&gt;
On a Belief MDP the agent is now able to choose between all possible actions to take as it may believe to be in any state with a given probability. We therefore must define a new variable &amp;lt;math&amp;gt;\pi &amp;lt;/math&amp;gt;that describes an action &amp;lt;math&amp;gt;a = \pi(b)&amp;lt;/math&amp;gt;given a certain belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The expected reward for the policy &amp;lt;math&amp;gt;\pi&lt;br /&gt;
&amp;lt;/math&amp;gt;starting from belief &amp;lt;math&amp;gt;b_0&amp;lt;/math&amp;gt;is &amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;&lt;br /&gt;
V^\pi(b_0) = \sum_{t=0}^\infty  \gamma^t r(b_t, a_t) = \sum_{t=0}^\infty \gamma^t E\Bigl[ R(s_t,a_t) \mid b_0, \pi \Bigr]&lt;br /&gt;
&amp;lt;/math&amp;gt;where again &amp;lt;math&amp;gt;\gamma &amp;lt; 1&amp;lt;/math&amp;gt;is the discount factor.&lt;br /&gt;
&lt;br /&gt;
By optimizing for long-term reward we can obtain the optimal policy &amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\pi ^{*}={\underset  {\pi }{{\mbox{argmax}}}}\ V^{\pi }(b_{0})&amp;lt;/math&amp;gt;where &amp;lt;math&amp;gt;b_0&amp;lt;/math&amp;gt;is the initial state.&lt;br /&gt;
&lt;br /&gt;
The optimal value function therefore can be described as&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;{\displaystyle V^{*}(b)=\max _{a\in A}{\Bigl [}r(b,a)+\gamma \sum _{o\in \Omega }\Pr(o\mid b,a)V^{*}(\tau (b,a,o)){\Bigr ]}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The value function is piecewise linear and convex for a finite-horizon POMDP&amp;lt;ref&amp;gt;Smallwood, R. D., &amp;amp; Sondik, E. J. (1973). &#039;&#039;The Optimal Control of Partially Observable Markov Processes over a Finite Horizon. Operations Research, 21(5), 1071–1088.&#039;&#039; doi:10.1287/opre.21.5.1071 &amp;lt;/ref&amp;gt;. For an infinite-horizon POMDP a finite vector set can approximate &amp;lt;math&amp;gt;V^*&amp;lt;/math&amp;gt;with the use of dynamic programming techniques.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the Markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[2] Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[3] Smallwood, R. D., &amp;amp; Sondik, E. J. (1973). &#039;&#039;The Optimal Control of Partially Observable Markov Processes over a Finite Horizon. Operations Research, 21(5), 1071–1088.&#039;&#039; doi:10.1287/opre.21.5.1071 &lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=586549</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=586549"/>
		<updated>2020-03-13T20:50:38Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2019|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]] (2020)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
* F1 [[Course:CPSC 522/Progressive Neural Network|Progressive Neural Network]] (2020)&lt;br /&gt;
* F5 [[Course:CPSC522/Conditional GANs for Image to Image Translation|Conditional GANs for Image-To-Image Translation]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* F2 [[Course:CPSC522/Variation Auto-Encoders |Variational Auto-Encoders]] (2020)&lt;br /&gt;
* F3[[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* F4 [[Course:CPSC522/Online Pattern Analysis by Evolving Self-Organizing Maps|Online Pattern Analysis by Evolving Self-organizing Maps]] (2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586547</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586547"/>
		<updated>2020-03-13T20:48:54Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)|623x623px]]&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment&amp;lt;ref&amp;gt;Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
*&amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
After executing one time step an agent needs to update its belief of the state of the environment it is in&amp;lt;ref name=&amp;quot;:0&amp;quot;&amp;gt;Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. The belief update is an actionless step and is therefore the same as seen on a [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Model] belief update. &lt;br /&gt;
&lt;br /&gt;
The belief state defined as &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; is a function over all states &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; from [0,1] that sums up to 1. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
[[File:DroneSurveilance.gif|thumb|Drone Navigating around space avoiding ground agent using POMDP|439x439px]]&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
An example of a POMDP in action can be seen on the left, this animation illustrates a drone (in yellow) navigating a tiled environment. Its objective is to reach the oposing green square without coming into contact with the ground agent (in red). The blue squares represent the drone&#039;s observable environment. Since the drone is not capale of observing the state of the ground agent at all times, MDP is not tractable for this task, POMDP enables it to estimate a probability distribution of the likely states the ground agent might be in and act accordingly.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the original POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
[[File:POMDP policy.png|thumb|POMDP decomposed into a state estimator and a policy.]]&lt;br /&gt;
&lt;br /&gt;
=== Policy and Value Function ===&lt;br /&gt;
On a Belief MDP the agent is now able to choose between all possible actions to take as it may believe to be in any state with a given probability. We therefore must define a new variable &amp;lt;math&amp;gt;\pi &amp;lt;/math&amp;gt;that describes an action &amp;lt;math&amp;gt;a = \pi(b)&amp;lt;/math&amp;gt;given a certain belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The expected reward for the policy &amp;lt;math&amp;gt;\pi&lt;br /&gt;
&amp;lt;/math&amp;gt;starting from belief &amp;lt;math&amp;gt;b_0&amp;lt;/math&amp;gt;is &amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;&lt;br /&gt;
V^\pi(b_0) = \sum_{t=0}^\infty  \gamma^t r(b_t, a_t) = \sum_{t=0}^\infty \gamma^t E\Bigl[ R(s_t,a_t) \mid b_0, \pi \Bigr]&lt;br /&gt;
&amp;lt;/math&amp;gt;where again &amp;lt;math&amp;gt;\gamma &amp;lt; 1&amp;lt;/math&amp;gt;is the discount factor.&lt;br /&gt;
&lt;br /&gt;
By optimizing for long-term reward we can obtain the optimal policy &amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\pi ^{*}={\underset  {\pi }{{\mbox{argmax}}}}\ V^{\pi }(b_{0})&amp;lt;/math&amp;gt;where &amp;lt;math&amp;gt;b_0&amp;lt;/math&amp;gt;is the initial state.&lt;br /&gt;
&lt;br /&gt;
The optimal value function therefore can be described as&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;{\displaystyle V^{*}(b)=\max _{a\in A}{\Bigl [}r(b,a)+\gamma \sum _{o\in \Omega }\Pr(o\mid b,a)V^{*}(\tau (b,a,o)){\Bigr ]}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The value function is piecewise linear and convex for a finite-horizon POMDP&amp;lt;ref&amp;gt;Smallwood, R. D., &amp;amp; Sondik, E. J. (1973). &#039;&#039;The Optimal Control of Partially Observable Markov Processes over a Finite Horizon. Operations Research, 21(5), 1071–1088.&#039;&#039; doi:10.1287/opre.21.5.1071 &amp;lt;/ref&amp;gt;. For an infinite-horizon POMDP a finite vector set can approximate &amp;lt;math&amp;gt;V^*&amp;lt;/math&amp;gt;with the use of dynamic programming techniques.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[2] Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[3] Smallwood, R. D., &amp;amp; Sondik, E. J. (1973). &#039;&#039;The Optimal Control of Partially Observable Markov Processes over a Finite Horizon. Operations Research, 21(5), 1071–1088.&#039;&#039; doi:10.1287/opre.21.5.1071 &lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:POMDP_policy.png&amp;diff=586543</id>
		<title>File:POMDP policy.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:POMDP_policy.png&amp;diff=586543"/>
		<updated>2020-03-13T20:29:07Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=POMDP decomposed into a state estimator and a policy.}}&lt;br /&gt;
|date=2020-03-13&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;br /&gt;
&lt;br /&gt;
[[Category:Computer Science]]&lt;br /&gt;
[[Category:Computers and Technology]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586530</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586530"/>
		<updated>2020-03-13T19:54:06Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)|623x623px]]&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment&amp;lt;ref&amp;gt;Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
*&amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
After executing one time step an agent needs to update its belief of the state of the environment it is in&amp;lt;ref&amp;gt;Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The belief state defined as &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; is a function over all states &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; from [0,1] that sums up to 1. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
[[File:DroneSurveilance.gif|thumb|Drone Navigating around space avoiding ground agent using POMDP|439x439px]]&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
An example of a POMDP in action can be seen on the left, this animation illustrates a drone (in yellow) navigating a tiled environment. Its objective is to reach the oposing green square without coming into contact with the ground agent (in red). The blue squares represent the drone&#039;s observable environment. Since the drone is not capale of observing the state of the ground agent at all times, MDP is not tractable for this task, POMDP enables it to estimate a probability distribution of the likely states the ground agent might be in and act accordingly.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the original POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[2] Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586525</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586525"/>
		<updated>2020-03-13T19:40:36Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)|623x623px]]&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
*&amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
after executing one time step an agent needs to update its belief of the state of the environment it is in. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
[[File:DroneSurveilance.gif|thumb|Drone Navigating around space avoiding ground agent using POMDP|443x443px]]&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
An example of a POMDP in action can be seen on the left, this animation illustrates a drone (in yellow) navigating a tiled environment. Its objective is to reach the oposing green square without coming into contact with the ground agent (in red). The blue squares represent the drone&#039;s observable environment. Since the drone is not capale of observing the state of the ground agent at all times, MDP is not tractable for this task, POMDP enables it to estimate a probability distribution of the likely states the ground agent might be in and act accordingly.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will thus be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the oroginal POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math display=&amp;quot;block&amp;quot;&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Åström, K.J. “Optimal Control of Markov Processes with Incomplete State Information.” &#039;&#039;Journal of Mathematical Analysis and Applications&#039;&#039; 10, no. 1 (February 1965): 174–205. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/0022-247X(65)90154-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[2] Kaelbling, Leslie Pack, Michael L. Littman, and Anthony R. Cassandra. “Planning and Acting in Partially Observable Stochastic Domains.” &#039;&#039;Artificial Intelligence&#039;&#039; 101, no. 1–2 (May 1998): 99–134. &amp;lt;nowiki&amp;gt;https://doi.org/10.1016/S0004-3702(98)00023-X&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586521</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586521"/>
		<updated>2020-03-13T19:23:20Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
[[File:DroneSurveilance.gif|thumb|Drone Navigating around space avoiding ground agent using POMDP]]&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)]]&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
* &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
after executing one time step an agent needs to update its belief of the state of the environment it is in. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will thus be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the oroginal POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows &amp;lt;math&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
[2]G. Hollinger, S. Singh, J. Djugash, and A. Kehagias, “Efficient Multi-robot Search for a Moving Target,” The International Journal of Robotics Research, vol. 28, no. 2, pp. 201–219, Feb. 2009, doi: 10.1177/0278364908099853.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:DroneSurveilance.gif&amp;diff=586511</id>
		<title>File:DroneSurveilance.gif</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:DroneSurveilance.gif&amp;diff=586511"/>
		<updated>2020-03-13T19:10:41Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Drone Navigating around space avoiding ground agent using POMDP}}&lt;br /&gt;
|date=2020-03-13&lt;br /&gt;
|source=https://juliaobserver.com/packages/POMDPGallery&lt;br /&gt;
|author= M. Svoreňová, M. Chmelík, K. Leahy, H. F. Eniser, K. Chatterjee, I. Černá, C. Belta&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{cc-by-sa-4.0}}&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586502</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=586502"/>
		<updated>2020-03-13T18:50:43Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs and [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]].&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
[[File:POMDP.io.png|thumb|Partially Obsevable Markov Decision Process example (POMDP)]]&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|Markov Decision Process&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|Hidden Markov Model&lt;br /&gt;
|Partilly Observable Markov Decision Process&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
* &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set that specifies the conditional probability of the next state givent he previous state and action&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\infty\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
after executing one time step an agent needs to update its belief of the state of the environment it is in. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will thus be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the oroginal POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows &amp;lt;math&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
[2]G. Hollinger, S. Singh, J. Djugash, and A. Kehagias, “Efficient Multi-robot Search for a Moving Target,” The International Journal of Robotics Research, vol. 28, no. 2, pp. 201–219, Feb. 2009, doi: 10.1177/0278364908099853.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:POMDP.io.png&amp;diff=586498</id>
		<title>File:POMDP.io.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:POMDP.io.png&amp;diff=586498"/>
		<updated>2020-03-13T18:39:49Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=Partially Obsevable Markov Decision Process example (POMDP)}}&lt;br /&gt;
|date=2020-03-13&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|TommasoDAmico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;br /&gt;
&lt;br /&gt;
[[Category:Computer Science]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Conditional_GANs_for_Image_to_Image_Translation/peer_feedback&amp;diff=586018</id>
		<title>Thread:Course talk:CPSC522/Conditional GANs for Image to Image Translation/peer feedback</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Conditional_GANs_for_Image_to_Image_Translation/peer_feedback&amp;diff=586018"/>
		<updated>2020-03-10T01:59:51Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;I really enjoyed reading about this topic! &lt;br /&gt;
&lt;br /&gt;
The examples are quite nice to show its capabilities,&lt;br /&gt;
&lt;br /&gt;
It would make it a bit clearer if you wrote the variables in latex when describing the objective function.&lt;br /&gt;
&lt;br /&gt;
I would like to perhaps know a bit more about the advantages and disadvantages of this algorithm, I found the last few paragraphs a bit confusing in their structure.&lt;br /&gt;
&lt;br /&gt;
(5) The topic is relevant for the course.&lt;br /&gt;
&lt;br /&gt;
(4) The writing is clear and the English is good.&lt;br /&gt;
&lt;br /&gt;
(5) The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds).&lt;br /&gt;
&lt;br /&gt;
(4) The formalism (definitions, mathematics) was well chosen to make the page easier to understand.&lt;br /&gt;
&lt;br /&gt;
(5) The abstract is a concise and clear summary.&lt;br /&gt;
&lt;br /&gt;
(5) There were appropriate (original) examples that helped make the topic clear.&lt;br /&gt;
&lt;br /&gt;
(-) There was appropriate use of (pseudo-) code.&lt;br /&gt;
&lt;br /&gt;
(4) It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic).&lt;br /&gt;
&lt;br /&gt;
(5) It is correct.&lt;br /&gt;
&lt;br /&gt;
(5) It was neither too short nor too long for the topic.&lt;br /&gt;
&lt;br /&gt;
(5) It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page).&lt;br /&gt;
&lt;br /&gt;
(5) It links to appropriate other pages in the wiki.&lt;br /&gt;
&lt;br /&gt;
(5) The references and links to external pages are well chosen.&lt;br /&gt;
&lt;br /&gt;
(5) I would recommend this page to someone who wanted to find out about the topic.&lt;br /&gt;
&lt;br /&gt;
(4) This page should be highlighted as an exemplary page for others to emulate.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Conditional_GANs_for_Image_to_Image_Translation/peer_feedback&amp;diff=586017</id>
		<title>Thread:Course talk:CPSC522/Conditional GANs for Image to Image Translation/peer feedback</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC522/Conditional_GANs_for_Image_to_Image_Translation/peer_feedback&amp;diff=586017"/>
		<updated>2020-03-10T01:59:13Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: New thread: peer feedback&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;I really enjoyed reading about this topic! &lt;br /&gt;
The examples are quite nice to show its capabilities,&lt;br /&gt;
It would make it a bit clearer if you wrote the variables in latex when describing the objective function.&lt;br /&gt;
I would like to perhaps know a bit more about the advantages and disadvantages of this algorithm, I found the last few paragraphs a bit confusing in their structure.&lt;br /&gt;
&lt;br /&gt;
(5) The topic is relevant for the course.&lt;br /&gt;
&lt;br /&gt;
(4) The writing is clear and the English is good.&lt;br /&gt;
&lt;br /&gt;
(5) The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds).&lt;br /&gt;
&lt;br /&gt;
(4) The formalism (definitions, mathematics) was well chosen to make the page easier to understand.&lt;br /&gt;
&lt;br /&gt;
(5) The abstract is a concise and clear summary.&lt;br /&gt;
&lt;br /&gt;
(5) There were appropriate (original) examples that helped make the topic clear.&lt;br /&gt;
&lt;br /&gt;
(-) There was appropriate use of (pseudo-) code.&lt;br /&gt;
&lt;br /&gt;
(4) It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic).&lt;br /&gt;
&lt;br /&gt;
(5) It is correct.&lt;br /&gt;
&lt;br /&gt;
(5) It was neither too short nor too long for the topic.&lt;br /&gt;
&lt;br /&gt;
(5) It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page).&lt;br /&gt;
&lt;br /&gt;
(5) It links to appropriate other pages in the wiki.&lt;br /&gt;
&lt;br /&gt;
(5) The references and links to external pages are well chosen.&lt;br /&gt;
&lt;br /&gt;
(5) I would recommend this page to someone who wanted to find out about the topic.&lt;br /&gt;
&lt;br /&gt;
(4) This page should be highlighted as an exemplary page for others to emulate.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC_522/Progressive_Neural_Network/Peer_Feedback&amp;diff=586013</id>
		<title>Thread:Course talk:CPSC 522/Progressive Neural Network/Peer Feedback</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Thread:Course_talk:CPSC_522/Progressive_Neural_Network/Peer_Feedback&amp;diff=586013"/>
		<updated>2020-03-10T01:25:17Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: New thread: Peer Feedback&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Your page is well structured and the concepts presented are well introduced. I&#039;d rephrase some parts of the introduction to make it easier to read. &lt;br /&gt;
for example &amp;quot;This paper will give a short overview of continual learning and the different approaches used in literature to achieve it. The paper will then explore one of these methods, progressive neural networks.&amp;quot;&lt;br /&gt;
maybe something of this kind: &amp;quot;This paper will provide a short overview of continual learning and the different approaches used, diving more into detail in the progressive neural networks method.&amp;quot;&lt;br /&gt;
It seems as if you are yet to finish your page. I&#039;d like to maybe have some examples of applications of the algorithm or where it is been used.&lt;br /&gt;
I liked your graph, it was very relevant and helpful to my comprehension.&lt;br /&gt;
&lt;br /&gt;
(5) The topic is relevant for the course.&lt;br /&gt;
&lt;br /&gt;
(4) The writing is clear and the English is good.&lt;br /&gt;
&lt;br /&gt;
(5) The page is written at an appropriate level for CPSC 522 students (where the students have diverse backgrounds).&lt;br /&gt;
&lt;br /&gt;
(4) The formalism (definitions, mathematics) was well chosen to make the page easier to understand.&lt;br /&gt;
&lt;br /&gt;
(5) The abstract is a concise and clear summary.&lt;br /&gt;
&lt;br /&gt;
(-) There were appropriate (original) examples that helped make the topic clear.&lt;br /&gt;
&lt;br /&gt;
(-) There was appropriate use of (pseudo-) code.&lt;br /&gt;
&lt;br /&gt;
(4) It had a good coverage of representations, semantics, inference and learning (as appropriate for the topic).&lt;br /&gt;
&lt;br /&gt;
(5) It is correct.&lt;br /&gt;
&lt;br /&gt;
(5) It was neither too short nor too long for the topic.&lt;br /&gt;
&lt;br /&gt;
(5) It was an appropriate unit for a page (it shouldn&#039;t be split into different topics or merged with another page).&lt;br /&gt;
&lt;br /&gt;
(5) It links to appropriate other pages in the wiki.&lt;br /&gt;
&lt;br /&gt;
(5) The references and links to external pages are well chosen.&lt;br /&gt;
&lt;br /&gt;
(5) I would recommend this page to someone who wanted to find out about the topic.&lt;br /&gt;
&lt;br /&gt;
(4) This page should be highlighted as an exemplary page for others to emulate.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584254</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584254"/>
		<updated>2020-03-02T21:37:22Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs, [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]] and game logic with Partially Observable Stochastic Games (POSG)&amp;lt;ref&amp;gt;Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
[[File:State Machine illustration.jpg|thumb|POMDP Example with two states s1,s2 and two actions a1 and a2]]&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|MDP&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|HMM&lt;br /&gt;
|POMDPs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
* &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of conditional transition probabilities between states&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\inf\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
after executing one time step an agent needs to update its belief of the state of the environment it is in. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will thus be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the oroginal POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows &amp;lt;math&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
==== Games ====&lt;br /&gt;
Some games (like poker) have hidden states, with POMDPs we can potentially compute a best response to a fixed opponent policy. Although solving the full game is a Partially Observable Stochastic Game (POSG) which is harder to solve than POMDP and is beyond the scope of this page.&lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
[2]G. Hollinger, S. Singh, J. Djugash, and A. Kehagias, “Efficient Multi-robot Search for a Moving Target,” The International Journal of Robotics Research, vol. 28, no. 2, pp. 201–219, Feb. 2009, doi: 10.1177/0278364908099853.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=584253</id>
		<title>Course:CPSC522</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522&amp;diff=584253"/>
		<updated>2020-03-02T20:58:21Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--Begin Infobox; Please add your parameters after the equal signs below.  If you do not wish to use the infobox, you may remove it by deleting everything between the Begin and End Infobox lines--&amp;gt;&lt;br /&gt;
{{Infobox_New_Course&lt;br /&gt;
&lt;br /&gt;
|title=CPSC 522 Wiki&lt;br /&gt;
&lt;br /&gt;
|picture=Image:wiki.png&lt;br /&gt;
&lt;br /&gt;
|subject code=CPSC&lt;br /&gt;
&lt;br /&gt;
|course number=522&lt;br /&gt;
&lt;br /&gt;
|instructor=David Poole&lt;br /&gt;
&lt;br /&gt;
|email=poole@cs.ubc.ca&lt;br /&gt;
&lt;br /&gt;
|office= 109&lt;br /&gt;
|office hours= after class every day&lt;br /&gt;
|classroom= DMP 101&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&amp;lt;!--End Infobox; Please add your page content below--&amp;gt;&lt;br /&gt;
[[Category:CPSC522]]&lt;br /&gt;
Welcome to [http://www.cs.ubc.ca/~poole/cs522/2019 CPSC 522] Wiki. This is where the participants are writing the textbook. See &lt;br /&gt;
http://www.cs.ubc.ca/~poole/cs522/2020/ for the main web page for the course.&lt;br /&gt;
==The 2020 Rules==&lt;br /&gt;
* These rules are editable, so you can change the rules.&lt;br /&gt;
* [[Course:CPSC522/StudentPresentations2020|2020 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 2 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2020|January and February Assignment]]&amp;lt;nowiki/&amp;gt;s describes your assignments for January and February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2019|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
==Old (2018, 2019) Rules==&lt;br /&gt;
*[[Course:CPSC522/StudentPresentations2018|2018 Student Presentation Schedule]] is the schedule for presentations. &#039;&#039;&#039;Please sign up early. This is first-come first-choice.&#039;&#039;&#039; There is a maximum of 3 students per day.&lt;br /&gt;
* [[Course:CPSC522/January2018|January Assignment]] describes your assignment for January.&lt;br /&gt;
* [[Course:CPSC522/February2018|February Assignment]] describes your assignment for February.&lt;br /&gt;
* [[Course:CPSC522/MarchApril2018|March-April Assignment]] describes your 3rd assignment.&lt;br /&gt;
* Do not plagiarize; give references to all sources. Put quotes in quotes.&lt;br /&gt;
* All pages should follow the [[Course:CPSC522/Template|CPSC522 Template]]&lt;br /&gt;
* Give links, but make sure that the pages are readable (with sentences) without following the links.&lt;br /&gt;
* Link to all prerequisite pages (those the page builds on) in the appropriate place in the template. The prerequisite pages should not form a cycle. &lt;br /&gt;
* Link to all related pages; it should be easy for someone to determine that a page on a particular topic does not exist.&lt;br /&gt;
* The web pages created should all be in the [[Course:CPSC522]] hierarchy, but should not contain &amp;quot;Course:CPSC522&amp;quot; in the visible link. Hint: create a link to the page before the page exists, then you will be asked to create the page,&lt;br /&gt;
* If you want to abandon a page that others are relying on (e.g., you are the principal author and they are coauthors) you need to negotiate with them to make sure no one is disadvantaged.&lt;br /&gt;
* All pages should be in the [[Course:CPSC522/Index|Index]] and the table of contents below.&lt;br /&gt;
&lt;br /&gt;
== Guidelines ==&lt;br /&gt;
* Keep each page as simple as possible (but not simpler); if a page starts to get complicated, consider splitting it.&lt;br /&gt;
* Write pages for your peers; they should all be written for incoming graduate students, and only assume background knowledge that is common among such students.&lt;br /&gt;
* All pages should obey the [[Course:CPSC522/Conventions|Syntax Conventions]]. If there is a design decision that you need to make that may have non-local implications, add it to the conventions.&lt;br /&gt;
* It should use formalism and mathematics when (and only when) the formalism make the description clearer. Use the code tags for math, e.g., &amp;lt;math&amp;gt;P(h\mid e) = \frac{P(h\land e)}{P(e)}.&amp;lt;/math&amp;gt;  It is worth your while to learn [https://www.latex-project.org/ Latex] if you don&#039;t already know it. &lt;br /&gt;
* If there is a simple case, and a more general case, give the simple case first. Making things complicated is easy; keeping them simple is difficult and we should strive for simplicity. Any complication needs to be carefully motivated.&lt;br /&gt;
* Use the &amp;quot;discussion&amp;quot; tab&lt;br /&gt;
&lt;br /&gt;
Each Page should contain:&lt;br /&gt;
* A clear jargon-free description of what is going on. Keep jargon to a minimum.&lt;br /&gt;
* Motivating example(s) and, where appropriate, a simple pedagogical example (which may be different from the motivating examples) that is used to explain what is going on&lt;br /&gt;
* An argument of plausibility&lt;br /&gt;
* Evidence that it works &lt;br /&gt;
* Code and pseudo-code, where appropriate. This code should interact with other related code (e.g., [http://aipython.org AIFCA Python Distribution]) if possible.  The code should be as simple as possible to implement the techniques. Consider adding exercises as to what can be improved or made more general or bullet-proof. Use a &amp;lt;code&amp;gt;code block&amp;lt;/code&amp;gt; for (pseudo-)code (even multi-line code). You can also use the format in http://wiki.ubc.ca/Course:CPSC_320/Midterm_2_Reference_Sheet#Pseudocode (try both and see which better suits your needs).&lt;br /&gt;
&lt;br /&gt;
==Foundations==&lt;br /&gt;
Please add your page here and in the [[Course:CPSC522/Index|Index]]. &lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/MyTest|My Test]]&lt;br /&gt;
===General===&lt;br /&gt;
* [[Course:CPSC522/AGI|Artificial General Intelligence]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Swarm_Intelligence|Swarm Intelligence]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Control===&lt;br /&gt;
* [[Course:CPSC522/Control Theory|Control Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Hierarchical Control|Hierarchical Control]] (2016)&lt;br /&gt;
&lt;br /&gt;
===Probability and Graphical Models===&lt;br /&gt;
* [[Course:CPSC522/Probability|Probability]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Graphical Models|Graphical Models]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Networks|Bayesian Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov_Networks|Markov Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/WeightedModelCounting|Weighted Model Counting]](2019)&lt;br /&gt;
====Temporal Models====&lt;br /&gt;
* [[Course:CPSC522/Markov Chains|Markov Chains]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Hidden_Markov_Models|Hidden Markov Models]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Kalman_filter|Kalman filter]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Dynamic Bayesian Networks|Dynamic Bayesian Networks]] (2018)&lt;br /&gt;
====Inference====&lt;br /&gt;
* [[Course:CPSC522/Variable Elimination|Variable Elimination]] (2016)&lt;br /&gt;
* [[Course:CPSC522/MCMC|Markov Chain Monte Carlo]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Particle Filtering|Particle Filtering]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Treatment of Missing Data|Treatment of Missing Data]] (2019)&lt;br /&gt;
* [[Course:CPSC522/Bayesian Coresets|Bayesian Coresets]] (2019)&lt;br /&gt;
&lt;br /&gt;
====Causality====&lt;br /&gt;
* [[Course:CPSC522/Causality|Causality]] (2016)&lt;br /&gt;
====Representations of Conditional Probability====&lt;br /&gt;
* [[Course:CPSC522/Neural Network|Neural Network]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Recurrent Neural Networks|Recurrent Neural Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Decision_Trees|Decision Trees]] (2018)&lt;br /&gt;
====Learning====&lt;br /&gt;
* [[Course:CPSC522/Learning Probabilistic Models with Complete Data|Learning Probabilistic Models with Complete Data]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Support_Vector_Machines|Support Vector Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ensemble Learning|Ensemble Learning]] (2018)&lt;br /&gt;
* J1 [[Course:CPSC522/Principal_Component_Analysis|Principal Component Analysis (PCA)]] (2020)&lt;br /&gt;
* J2 [[Course:CPSC 522/Self-Organizing Maps|Self-Organizing Maps]] (2020)&lt;br /&gt;
===NLP===&lt;br /&gt;
* [[Course:CPSC522/Natural Language Processing | Natural Language Processing]] (2018)&lt;br /&gt;
* [[Course:CPSC522/PCFG|Probabilistic Context Free Grammars]] (2018)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Utility and Preferences===&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Bounded Rationality|Bounded Rationality]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Elicitation of Factored Utilities|Elicitation of Factored Utilities]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Acting Under Uncertainty===&lt;br /&gt;
* [[Course:CPSC522/Decision Networks|Decision Networks]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Markov Decision Process|Markov Decision Process]] (2016)&lt;br /&gt;
* F0 [[Course:CPSC522/Partially Observable Markov Decision Processes|Partially Observable Markov Decision Processes]]&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning|Reinforcement Learning]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning with Function Approximation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Game Theory|Game Theory]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Multi-Agent Systems|Multi-Agent Systems]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Stochastic Optimization|Stochastic Optimization]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Value of Information and Control|Value of Information]] (2019)&lt;br /&gt;
&lt;br /&gt;
===Logic===&lt;br /&gt;
*  [[Course:CPSC522/Abduction|Abduction]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Knowledge Compilation|Knowledge Compilation]] (2016)&lt;br /&gt;
* [[Course:CPSC522/Predicate Calculus|Predicate Calculus]] (2016)&lt;br /&gt;
*  [[Course:CPSC522/Markov Logic|Markov Logic]] (2018)&lt;br /&gt;
*  [[Course:CPSC522/Higher Order Logic|Higher Order Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Ontology|Ontology]] (2019)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Recommendation System using Matrix Factorization|Recommendation System using Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Latent Dirichlet Allocation|Latent Dirichlet Allocation]]&lt;br /&gt;
* [[Course:CPSC522/Deep Neural Network|Deep Neural Network and Game of Go]]&lt;br /&gt;
* [[Course:CPSC522/Problog|Problog]]&lt;br /&gt;
* [[Course:CPSC522/Maximum Entropy Markov Models|Maximum Entropy Markov Models]]&lt;br /&gt;
* [[Course:CPSC522/Future Directions for Semantic Systems|Ontology Search Engine]]&lt;br /&gt;
* [[Course:CPSC522/Convolutional Neural Networks|Convolutional Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Learning Markov Logic Network Structure|Learning Markov Logic Network Structure]]&lt;br /&gt;
* [[Course:CPSC522/Decision Support System using Interactive Preference Elicitation|Decision Support System using Interactive Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System|Predicting Affect of User&#039;s Interaction with an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Record Linkage and identity uncertainty|Record Linkage and identity uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Robot Scientist|Robot Scientist]]&lt;br /&gt;
* [[Course:CPSC522/Density-Based Unsupervised Learning|Density-Based Unsupervised Learning]]&lt;br /&gt;
* [[Course:CPSC522/Predicting Human Behavior in Normal-Form Games|Predicting Human Behavior in Normal-Form Games]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty|Identity Uncertainty]]&lt;br /&gt;
* [[Course:CPSC522/Generative Adversarial Networks|Generative Adversarial Networks]]&lt;br /&gt;
&amp;lt;!-- *[[Course:CPSC522/Ontology|Ontology]] Sorry for not removing this page earlier. Samprity had already taken the same topic --&amp;gt;&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
*[[Course:CPSC522/User-Adaptive Information Visualization|User-Adaptive Information Visualization]]&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2020)==&lt;br /&gt;
* J0 [[Course:CPSC522/Variational Inference|Variational Inference]]&lt;br /&gt;
* J3 [[Course:CPSC522/Monte Carlo Localization |Pedestrian localization for Indoor Environments]] (2020)&lt;br /&gt;
* J4 [[Course:CPSC522/Automation of hypothesis generation and testing in science|Automation of hypothesis generation and testing in science]] (2020)&lt;br /&gt;
* J5 [[Course:CPSC522/Deep Q Network |Prioritized Experience Replay]] (2020)&lt;br /&gt;
* [[Course:CPSC522/Combining_Collaborative_Filtering_with_Personal_Agents_for_Better_Recommendations | Hybrid Recommendation Systems]] (2020)&lt;br /&gt;
* [[On-line Pattern Analysis by Evolving Self-organizing Maps]](2020)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2018)==&lt;br /&gt;
===Neural Networks===&lt;br /&gt;
* [[Course:CPSC522/Financial Forecasting using LSTM Networks |Financial Forecasting using LSTM Networks]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Character Level Language Models using LSTM|Character Level Language Models using LSTM]] (2018)&lt;br /&gt;
* [[Course:CPSC522/TextSummarizationUsingMachineLearning |Text Summarization using Machine Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Password_cracking_using_PCFGs_and_Neural_Networks|Password Cracking using Probabilistic Context Free Grammars and Neural Networks]] (2018)&lt;br /&gt;
* [[CNNs in Image Segmentation]](2018)&lt;br /&gt;
* [[Course:CPSC522/Image_Classification_With_Convolutional_Neural_Networks|Image Classification With Convolutional Neural Networks]] (2018)&lt;br /&gt;
* [[Image Colourization using Deep Learning]](2018)&lt;br /&gt;
* [[Course:CPSC522/StackedGAN|Stacked Generative Adversarial Networks]] (2018)&lt;br /&gt;
&lt;br /&gt;
===Reinforcement Learning===&lt;br /&gt;
* [[Course:CPSC522/Deep_Q-Learning|Deep Q-Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Deep_Reinforcement_Learning|Deep Reinforcement Learning]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Self_Improving_Machines|Self-Improving Machines]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Adaptive_Network_Routing_using_ACO|Adaptive Network Routing using Ant Colony Optimization]] (2018)&lt;br /&gt;
===Decision-theoretic Planning===&lt;br /&gt;
* [[Course:CPSC522/Action_Selection_for_MDPs|Action Selection for MDPs]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Rao_Blackwellized_Particle_Filtering|Rao-Blackwellized Particle Filtering]](2018)&lt;br /&gt;
===Relational Reasoning===&lt;br /&gt;
* [[Course:CPSC522/Transfer_Learning_with_Markov_Logic|Transfer Learning with Markov Logic]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Cognitive_Robotics|Cognitive Robotics]] (2018)&lt;br /&gt;
===Applications===&lt;br /&gt;
* [[Course:CPSC522/Affect Prediction using Eye Gaze|Affect Prediction using Eye Gaze]] (2018)&lt;br /&gt;
* [[Course:CPSC522/Conflict-Driven Clause Learning for the Boolean Satisfiability Problem|Conflict-Driven Clause Learning for the Boolean Satisfiability Problem]] (2018)&lt;br /&gt;
&lt;br /&gt;
==Existing Combinations (2019)==&lt;br /&gt;
* [[Course:CPSC522/Ontology Extraction|Ontology Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Restricted Boltzmann Machines for Collaborative Filtering|Restricted Boltzmann Machines for Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/Minimax Regret Preference Elicitation for Risky Prospects|Minimax Regret Preference Elicitation for Risky Prospects]]&lt;br /&gt;
* [[Course:CPSC522/Sequential Monte Carlo samplers|Sequential Monte Carlo samplers]]&lt;br /&gt;
* [[Course:CPSC522/FastSLAM|FastSLAM]]&lt;br /&gt;
* [[Course:CPSC522/SMC for PGMs|Sequential Monte Carlo for Probabilistic Graphical Models]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2016)==&lt;br /&gt;
* [[Course:CPSC522/Sentiment Analysis|Sentiment Analysis: Movie Reviews]]&lt;br /&gt;
* [[Course:CPSC522/Collaborative Filtering|Collaborative Filtering]]&lt;br /&gt;
* [[Course:CPSC522/List Recommendation|List Recommendation]]&lt;br /&gt;
* [[Course:CPSC522/Inactive Cookie Mapping via Trail Matching|Inactive Cookie Mapping via Trail Matching]]&lt;br /&gt;
* [[Course:CPSC522/Improve recommendation system by integration|Improve Recommendation System by Integration]]&lt;br /&gt;
* [[Course:CPSC522/Identity Uncertainty in a restaurant data-set|Identity Uncertainty in a restaurant data-set]]&lt;br /&gt;
* [[Course:CPSC522/Regularization_for_Neural_Networks|Regularization for Neural Networks]]&lt;br /&gt;
* [[Course:CPSC522/Spam Detection|Spam Detection]]&lt;br /&gt;
* [[Course:CPSC522/Titanic: Machine Learning from Disaster|Titanic: Machine Learning from Disaster]]&lt;br /&gt;
* [[Course:CPSC522/Automatic Classification of Morphological Heart Arrhythmia | Automatic Classification of Morphological Heart Arrhythmia]]&lt;br /&gt;
* [[Course:CPSC522/Linking Sentences in Asynchronous Conversations|Linking Sentences in Asynchronous Conversations]]&lt;br /&gt;
* [[Course:CPSC522/Generic Aspect-based Aggregation of Sentiments|Generic Aspect-based Aggregation of Sentiments]]&lt;br /&gt;
* [[Course:CPSC522/Graph Based keyword extraction|Graph Based Key-corporation Extraction]]&lt;br /&gt;
* [[Course:CPSC522/Improving the accuracy of Affect Prediction in an Intelligent Tutoring System|Improving the accuracy of Affect Prediction in an Intelligent Tutoring System]]&lt;br /&gt;
* [[Course:CPSC522/Improving Human Behavior Prediction in Simultaneous-Move Games|Improving Human Behavior Prediction in Simultaneous-Move Games]]&lt;br /&gt;
* [[Course:CPSC522/The Automation of Disease Diagnosis|The Automation of Disease Diagnosis]]&lt;br /&gt;
* [[Course:CPSC522/Analyzing online dating trends with Weka|Analyzing online dating trends with Weka]]&lt;br /&gt;
&lt;br /&gt;
==Future combinations (2018)==&lt;br /&gt;
*  [[Course:CPSC522/Artificial Intelligence and Economic Theory|Artificial Intelligence and Economic Theory]]&lt;br /&gt;
*  [[Course:CPSC522/Weak Semantic Map|Weak Semantic Map: Simplified Chinese]]&lt;br /&gt;
* [[Course:CPSC522/Network Agent|Datacenter Traffic as Reinforcement Learning Problem]]&lt;br /&gt;
*  [[Course:CPSC522/Baseilne_of_RSI|A Theoretical Baseline of Recursive Self-improvement]]&lt;br /&gt;
*  [[Course:CPSC522/Text_Summarization_for_busy_people!| Text summarization for busy people!!]]&lt;br /&gt;
*  [[Course:CPSC522/An_evaluation_on_selecting_and_applying_Recommendation_Methods | An evaluation on selecting and applying Recommendation Methods]]&lt;br /&gt;
*  [[Course:CPSC522/Learning User Preferences of Motion Control | Learning User Preferences of Motion Control]]&lt;br /&gt;
*  [[Course:CPSC522/Experiments_with_Reinforcement_Learning| Experiments with Reinforcement Learning]]&lt;br /&gt;
*  [[Course:CPSC522/A_Comparison_of_LDA_and_NMF_for_Topic_Modeling_on_Literary_Themes| A Comparison of LDA and NMF for Topic Modeling on Literary Themes]]&lt;br /&gt;
*  [[Course:CPSC522/Analysis of hierarchical prior for Language modeling | Analysis of hierarchical prior for Language modeling]]&lt;br /&gt;
*  [[Better_caching_using_reinforcement_learning|Better Caching using reinforcement learning]]&lt;br /&gt;
*  [[Course:CPSC522/Evaluation_of_ACO|Evaluating Ant Colony Optimization in a simulation]]&lt;br /&gt;
*  [[Course:CPSC522/SLAM_And_Sensor_Quality|SLAM and Sensor Quality]]&lt;br /&gt;
*  [[Text generation with LSTM and Markov Chain]]&lt;br /&gt;
*  [[Course:CPSC522/Topology_and_Embedding_Multi-relational_Data|Topology and Embedding Multi-relational Data]]&lt;br /&gt;
&lt;br /&gt;
== Future combinations (2019) ==&lt;br /&gt;
* [[Course:CPSC522/Reinforcement Learning with Linear Model of Reward Corruption|Reinforcement Learning with Linear Model of Reward Corruption]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Adverserial Belief Propagation|Adversarial Belief Propagation]]  &lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Using Subset Information with Matrix Factorization|Using Subset Information with Matrix Factorization]]&lt;br /&gt;
*&lt;br /&gt;
* [[Course:CPSC522/Learning Attention via Active Inference|Learning Attention via Active Inference]] &lt;br /&gt;
* [[Course:CPSC522/Regularization as an Alternative to Negative Sampling in KGs|Regularization as an Alternative to Negative Sampling in KGs]] &lt;br /&gt;
==Suggested Unclaimed Pages==&lt;br /&gt;
Here are some possible topics for pages. This list is not meant to limit your imagination. Some of them might be better split into multiple pages. There are many other possible topics.&lt;br /&gt;
&lt;br /&gt;
When claimed, these pages should be moved from this section to the table of contents above and to the  [[Course:CPSC522/Index|Index]] of existing pages. To claim a page you have to actually create it and edit it (and have your name on the page, so everyone can see who has claimed it).&lt;br /&gt;
&lt;br /&gt;
* [[Course:CPSC522/Probability general semantics|Probability - general semantics]] with infinitely many variables and/or variables with infinite domains&lt;br /&gt;
* [[Course:CPSC522/Representations of Conditional Distributions|Representations of Conditional Distributions]]&lt;br /&gt;
* [[Course:CPSC522/Recursive Conditioning|Recursive Conditioning]]&lt;br /&gt;
* [[Course:CPSC522/Parity Methods|Parity Methods for Probabilistic Inference]]&lt;br /&gt;
* [[Course:CPSC522/Matrix Factorization|Matrix Factorization]]&lt;br /&gt;
* [[Course:CPSC522/Utility|Utility]]&lt;br /&gt;
* [[Course:CPSC522/Multi-Attribute Utility|Multi-Attribute Utility]]&lt;br /&gt;
* [[Course:CPSC522/Preference Elicitation|Preference Elicitation]]&lt;br /&gt;
* [[Course:CPSC522/Mechanism Design|Mechanism Design]]&lt;br /&gt;
* [[Course:CPSC522/Logic Programming|Logic Programming]] &lt;br /&gt;
* [[Course:CPSC522/Negation as Failure|Negation as Failure]]&lt;br /&gt;
* [[Course:CPSC522/Equality-Identity|Equality/Identity]]&lt;br /&gt;
* [[Course:CPSC522/Inductive Logic Programming|Inductive Logic Programming]]&lt;br /&gt;
* [[Course:CPSC522/Ontologies|Ontologies]]&lt;br /&gt;
* [[Course:CPSC522/Continual Learning|Continual Learning]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584217</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584217"/>
		<updated>2020-03-02T02:42:48Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: First Draft Complete&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the potential of POMDPs for belief MDP.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs, [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]] and game logic with Partially Observable Stochastic Games (POSG)&amp;lt;ref&amp;gt;Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
[[File:State Machine illustration.jpg|thumb|POMDP Example with two states s1,s2 and two actions a1 and a2]]&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{|&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|MDP&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|HMM&lt;br /&gt;
|POMDPs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
* &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of conditional transition probabilities between states&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\inf\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favoured over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
==== Belief Update ====&lt;br /&gt;
after executing one time step an agent needs to update its belief of the state of the environment it is in. Since we assume a &#039;&#039;Markovian&#039;&#039; state space, all we need to describe to execute this step is our prior state belief &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, the last action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;taken by the agent and the last observation made &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The update step can be denoted as &amp;lt;math&amp;gt;b&#039; = \tau(b,a,o)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After reaching state &amp;lt;math&amp;gt;s&#039;&amp;lt;/math&amp;gt;, the agent observes &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;as described in the prior section. &lt;br /&gt;
&lt;br /&gt;
If we then let &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; be a probability distribution over the state space &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;then &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;denotes the probability that the environment is in state &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. Given &amp;lt;math&amp;gt;b(s)&amp;lt;/math&amp;gt;, then after taking an action &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;and an observation &amp;lt;math&amp;gt;o&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;b&#039;(s&#039;) = \eta O(o|s&#039;,a) \sum_{s \in S} T(s&#039; |s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;\eta = 1 / Pr(o|b,a)&amp;lt;/math&amp;gt; is used as a normalising constant with &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(o|b,a) = \sum_{s&#039; \in S} O(o|s&#039;,a) \sum_{s \in S} T(s&#039; | s,a)b(s)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Interpretation of Definition ====&lt;br /&gt;
From the above definition the agent does not directly observe its environment. It instead has to make decisions about the true environment state under uncertainty. However, by interacting with the environment and receiving observations the agent is able to update its belief of the true state by updating the probability distribution of the current state. Through this property of the algorithm the extrapolated optimal behaviour may often include actions that are taken purely because they improve the agent&#039;s estimate of the current state, thereby allowing it to make better decisions in future time steps. &lt;br /&gt;
&lt;br /&gt;
If we compare the formal definition of the POMDP described above with that of MDP we would have a very similar structure with the exception that the MDP algorithm would not include the conditional observation probabilities set  &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;as it is always certain of the observed true state.&lt;br /&gt;
&lt;br /&gt;
=== Belief MDP ===&lt;br /&gt;
A Markovian belief state allows a POMDP to be formulated as an Markov Decision Process where every belief is a state. The resulting belief MDP will thus be defined on a continuous state space even though the the originating POMDP has a finite number of states. This is because there are infinite number of probability distributions over the state set.&lt;br /&gt;
&lt;br /&gt;
The belief MDP is formally described as a tuple &amp;lt;math&amp;gt;(B,A,\tau ,r ,\gamma )&amp;lt;/math&amp;gt;where&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;is the set of belief states over the POMDP states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;is the same finite set of actions as in the POMDP algorithm&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau&amp;lt;/math&amp;gt;is the belief state transition function&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;r: B \times A \rightarrow \R&amp;lt;/math&amp;gt;is the reward function on belief states&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt;is the discount factor equal to the one in the original POMDP&lt;br /&gt;
&lt;br /&gt;
from these values &amp;lt;math&amp;gt;\tau&lt;br /&gt;
&amp;lt;/math&amp;gt;and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;are derived from the oroginal POMDP via&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\tau (b,a,b&#039;) = \sum_{o \in \Omega} Pr(b&#039;|b,a,o)Pr(o|a,b)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where &amp;lt;math&amp;gt;Pr(o|a,b)&amp;lt;/math&amp;gt; is the value derived in the previous section and &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Pr(b&#039;|b,a,o) = \begin{cases} 1, &amp;amp; \text{if the belief update with arguments } b,a,o \text{ returns } b&#039; \\ 0, &amp;amp; \text{otherwise} \end{cases}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
and the reward function is the expected reward from the POMDP reward function over the belief distribution as follows &amp;lt;math&amp;gt;r(b,a) = \sum_{s \in S} b(s)R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By going through these steps the belief MDP is not partially observable anymore, since at any given time the agent  knows its belief, and by extension the state of the belief MDP. &lt;br /&gt;
&lt;br /&gt;
This means we were able to take a partially observable discrete space system and convert it to a continuous space fully observable belief system.&lt;br /&gt;
&lt;br /&gt;
=== Applications of POMDP ===&lt;br /&gt;
&lt;br /&gt;
==== Pursuit-Evasion ====&lt;br /&gt;
Imagine a fleet of robots in a constrained environment, assign one robot to pursuit the others. The pursuer&#039;s state is known, but the evader&#039;s state is only partially observed. POMDP can be applied by the evaders in a multi-agent search fashion to observe the environment and build a belief of the pursuer&#039;s location.&amp;lt;ref&amp;gt;{{Cite journal|last=Hollinger, Singh|first=|date=|title=Efficient Multi-robot Search for a Moving Target|url=https://journals.sagepub.com/doi/10.1177/0278364908099853|journal=The International Journal of Robotics Research|volume=|pages=|via=}}&amp;lt;/ref&amp;gt; &lt;br /&gt;
&lt;br /&gt;
==== Sensor Placement ====&lt;br /&gt;
A store may have a limited amount of cameras to allocate in its interior to monitor its environment, since the &amp;quot;world &amp;quot; will be partially observable, using POMDP can help reconstruct intruder position or in the contrary facilitate &amp;quot;stealthy&amp;quot; movement. &lt;br /&gt;
&lt;br /&gt;
Games&lt;br /&gt;
&lt;br /&gt;
Some games (like poker) have hidden states, with POMDPs we can potentially compute a best response to a fixed opponent policy. Although solving the full game is a Partially Observable Stochastic Game (POSG) which is harder to solve than POMDP and is beyond the scope of this page.&lt;br /&gt;
&lt;br /&gt;
=== Conclusion ===&lt;br /&gt;
Probability theory is a powerful tool for modelling action under uncertainty, thanks to the markov assumption many algorithms have become easy to implement for this purpose. POMDPs are one of the broader algorithms in this category as they take structure from highly constrained algorithms such as MDPs and relax assumptions of observability to allow for more real-world problem solving applications. &lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
[2]G. Hollinger, S. Singh, J. Djugash, and A. Kehagias, “Efficient Multi-robot Search for a Moving Target,” The International Journal of Robotics Research, vol. 28, no. 2, pp. 201–219, Feb. 2009, doi: 10.1177/0278364908099853.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=File:State_Machine_illustration.jpg&amp;diff=584216</id>
		<title>File:State Machine illustration.jpg</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=File:State_Machine_illustration.jpg&amp;diff=584216"/>
		<updated>2020-03-02T01:52:21Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: User created page with UploadWizard&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=={{int:filedesc}}==&lt;br /&gt;
{{Information&lt;br /&gt;
|description={{en|1=State Machine Illustration with 2 States and 2 actions}}&lt;br /&gt;
|date=2020-03-01&lt;br /&gt;
|source={{own}}&lt;br /&gt;
|author=[[User:TommasoDAmico|Tommaso D&#039;Amico]]&lt;br /&gt;
|permission=&lt;br /&gt;
|other versions=&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=={{int:license-header}}==&lt;br /&gt;
{{self|cc-by-sa-4.0}}&lt;br /&gt;
&lt;br /&gt;
[[Category:CPSC522]]&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584215</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584215"/>
		<updated>2020-03-01T23:25:58Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: added formal definition&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the differences between POMDPs and other Markov Processes.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs, [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]] and game logic with Partially Observable Stochastic Games (POSG)&amp;lt;ref&amp;gt;Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
POMDPs are also closely related to [https://en.wikipedia.org/wiki/Hidden_Markov_model Hidden Markov Models] and [https://en.wikipedia.org/wiki/Markov_chain Markov Chains].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Introduction ===&lt;br /&gt;
Partially Observable Markov Decision Processes (POMDPs) are a type of Markov Process closely related to Markov Decision Processes (MDP) as:&lt;br /&gt;
* there exist a finite number of discrete states&lt;br /&gt;
* the next state is only determined by the current state and the current action taken by the agent&lt;br /&gt;
* there exists a probabilistic transition between states and controllable actions in each state&lt;br /&gt;
The way a POMDP differs from an MDP is that the agent is unsure of which state it is in as it has only partial observability of the environment states. It instead needs to develop a probability distribution over the set of states it believes to be in and condition its observations and the underlying MDP structure. A comparison can be drawn for clarity to the relationship of Markov Chains with Hidden Markov Models, as the latter too is an extension of the prior with reduced observability. In fact a helpful illustration of the difference among these models can be seen in the chart below:&lt;br /&gt;
{|&lt;br /&gt;
|+&#039;&#039;&#039;Helpful Chart for Markov Model Segmentation&#039;&#039;&#039;&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;2&amp;quot; |&lt;br /&gt;
| colspan=&amp;quot;2&amp;quot; rowspan=&amp;quot;1&amp;quot; |&#039;&#039;&#039;Does the agent have control over state transition?&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|yes&lt;br /&gt;
|-&lt;br /&gt;
| colspan=&amp;quot;1&amp;quot; rowspan=&amp;quot;2&amp;quot; |&#039;&#039;&#039;Are the states fully observable?&#039;&#039;&#039;&lt;br /&gt;
|yes&lt;br /&gt;
|Markov Chain&lt;br /&gt;
|MDP&lt;br /&gt;
|-&lt;br /&gt;
|no&lt;br /&gt;
|HMM&lt;br /&gt;
|POMDPs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== POMDP versus MPD ====&lt;br /&gt;
MPD is more tractable to solve and is relatively easy to specify although it assumes perfect knowledge of the states which is often unlikely for many use cases.&lt;br /&gt;
&lt;br /&gt;
POMDP on the other hand treats all sources of uncertainty uniformly and allows for information gathering actions, although it is hugely intractable to solve optimally.&lt;br /&gt;
&lt;br /&gt;
=== Formal Definition ===&lt;br /&gt;
Formally a POMDP working in discrete-time models the relationship between an agent and its environment. This is commonly done with a 7-tuple &amp;lt;math&amp;gt;(S,A,T,R,\Omega, O,\gamma)&amp;lt;/math&amp;gt;where:&lt;br /&gt;
* &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;is a set of all states&lt;br /&gt;
* &amp;lt;math&amp;gt;A&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of all actions&lt;br /&gt;
* &amp;lt;math&amp;gt;T&lt;br /&gt;
&amp;lt;/math&amp;gt;is a set of conditional transition probabilities between states&lt;br /&gt;
* &amp;lt;math&amp;gt;R:S \times A\rightarrow \R  &lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward function&lt;br /&gt;
* &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;is the set of observations&lt;br /&gt;
* &amp;lt;math&amp;gt;O&amp;lt;/math&amp;gt;is the set of conditional observation probabilities &lt;br /&gt;
* &amp;lt;math&amp;gt;\gamma \in [0,1]&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor&lt;br /&gt;
At each time step the environment is in state &amp;lt;math&amp;gt;s \in S&lt;br /&gt;
&amp;lt;/math&amp;gt;. The agent takes action &amp;lt;math&amp;gt;a \in A&lt;br /&gt;
&amp;lt;/math&amp;gt;, which causes the environment to transition to &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&amp;lt;/math&amp;gt;with probability given by &amp;lt;math&amp;gt;T(s&#039; | s,a).&lt;br /&gt;
&amp;lt;/math&amp;gt;at the same time the agent receives an observation &amp;lt;math&amp;gt;o \in \Omega&lt;br /&gt;
&amp;lt;/math&amp;gt;which is dependent on the new environment state &amp;lt;math&amp;gt;s&#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;and the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt;just done by the agent, this is described by the probability distribution of &amp;lt;math&amp;gt;O(o|s&#039;,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This state and action is then used to develop a reward &amp;lt;math&amp;gt;r&lt;br /&gt;
&amp;lt;/math&amp;gt;equal to &amp;lt;math&amp;gt;R(s,a)&lt;br /&gt;
&amp;lt;/math&amp;gt;. This process then repeats for the next time step until the agent completes its task or to infinity. The goal is for the agent to maximize its expected future discounted reward: &amp;lt;math&amp;gt;E[ \sum_{t=0}^\inf\gamma^tr_t ]&lt;br /&gt;
&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;r_t&lt;br /&gt;
&amp;lt;/math&amp;gt;is the reward earned at time &amp;lt;math&amp;gt;t&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;\gamma&lt;br /&gt;
&amp;lt;/math&amp;gt;is the discount factor that determines how much immediate rewards are favored over distant rewards. When &amp;lt;math&amp;gt;\gamma = 0&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent only cares about which action will yield the largest expected immediate reward, while if &amp;lt;math&amp;gt;\gamma = 1&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;the agent cares about maximizing the expected sum of future rewards.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584211</id>
		<title>Course:CPSC522/Partially Observable Markov Decision Processes</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Partially_Observable_Markov_Decision_Processes&amp;diff=584211"/>
		<updated>2020-03-01T22:21:12Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: Setup page for POMDPs&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partially Observable Markov Decision Processes (POMDPs) ==&lt;br /&gt;
A Partially Observable Markov Decision Processes (POMDs) is a mathematical model for acting under uncertainty. It expands from the Markov Decision Process (MDP) by relaxing the constraint of having a fully observable state space. In this paper we will evaluate the underlying mathematical model and explore the differences between POMDPs and other Markov Processes.&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators:&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
POMDPs are a type of Decision Network that expands from Markov Decision Processes by not requiring all states to be observable by the agent. POMDPs still maintain the same system dynamics of an MDP although the acting agent cannot always observe its current state but rather it needs to maintain a probability distribution over the set of possible states it may be in based on its observations and the underlying MPD.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
POMDPs is a mathematical model expanding from the concepts of [https://wiki.ubc.ca/Course:CPSC522/Decision_Networks Decision networks], specifically generalizing from the [[Course:CPSC522/Markov Decision Processes|Markov Decision Process]] algorithm for acting under uncertainty. &lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
POMDPs are widely implemented in applications that interact with the real world, some examples are particle filtering techniques such as MC-POMDPs, an extension of [https://wiki.ubc.ca/Course:CPSC522/MCMC Markov Chains Monte Carlo] for POMDPs, [[Course:CPSC522/Reinforcement Learning with Function approximation|Reinforcement Learning]] and game logic with Partially Observable Stochastic Games (POSG)&amp;lt;ref&amp;gt;Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
Put the content here. Use appropriate subheadings and links.&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Hansen, Eric A, Daniel S Bernstein, and Shlomo Zilberstein. “Dynamic Programming for Partially Observable Stochastic Games,” n.d., 6.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Monte_Carlo_Localization&amp;diff=582425</id>
		<title>Course:CPSC522/Monte Carlo Localization</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Monte_Carlo_Localization&amp;diff=582425"/>
		<updated>2020-02-10T21:29:44Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Monte Carlo Pedestrian Localization for Indoor Environments ==&lt;br /&gt;
This page gives an overview of the Monte Carlo Localization method (MCL)&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; and demonstrates an extension of the technique to enable for pedestrian tracking within a building environment.&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
Papers Discussed:&lt;br /&gt;
&lt;br /&gt;
- Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&lt;br /&gt;
&lt;br /&gt;
- Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
This page gives an overview of the Monte Carlo Localization method (MCL)&amp;lt;ref name=&amp;quot;:0&amp;quot;&amp;gt;Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&amp;lt;/ref&amp;gt; and demonstrates an extension of the technique to enable for pedestrian tracking within a building environment&amp;lt;ref name=&amp;quot;:1&amp;quot;&amp;gt;Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. MCL is analyzed as a technique for  pedestrian localization as it allows for global localization with multi modal probability distributions over continuous space and is shown to be robust to highly noisy sensors. The Pedestrian Localization paper demonstrates a technique to extend MCL to a 2.5D map and illustrates the benefits of it for reducing uncertainty.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
The Monte Carlo Localization method is an algorithm that utilizes [[Course:CPSC522/Particle Filtering|particle filtering]] as a means of localization. It builds on the idea of using [[Course:CPSC522/Probability|Bayesian probability]] with [[Course:CPSC522/Markov Chains|Markov Chains]] to develop a [[Course:CPSC522/Probability|probability distribution]] of the likely states of a robot to aid it in its localization. &lt;br /&gt;
&lt;br /&gt;
For the adaptation of MCL to pedestrians in indoors environments the [https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence Kullback-Leibler divergence (KLD)] sampling algorithm is used to dynamically adapt the number of particles sampled.&lt;br /&gt;
&lt;br /&gt;
=== Related Pages ===&lt;br /&gt;
a related combination of Monte Carlo Method with Markov Chains is the [[Course:CPSC522/MCMC|MCMC algorithm]].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Monte Carlo Localization (MCL) ===&lt;br /&gt;
The goal of Monte Carlo Localization is to allow a robot to localize itself in an environment given only an odometry system, a sensor to perceive its surroundings and a 2D map of its navigable environment&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In this page we will also explore the use of MCL without the aid of a sensor to perceive the surroundings&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt; and how it is still capable of providing accurate estimates using different aids and assumptions.&lt;br /&gt;
&lt;br /&gt;
==== Global Localization and Position Tracking ====&lt;br /&gt;
Localization for robotics is often broken down in two categories: Global localization and position tracking&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In position tracking the robot knows its initial position and only needs to accommodate for small errors or drift in the odometry system as it moves. Global Localization is a bit more complex as it looks at identifying the robots&#039; position within its global environment given no initial state.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; This is also referred to as the [https://en.wikipedia.org/wiki/Kidnapped_robot_problem hijacked robot problem]. Since the robot needs to localize itself from scratch, this tends to be a very computationally heavy process and is often omitted by most localization algorithms.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====  MCL Algorithm ====&lt;br /&gt;
[[File:Sca80a0 Animation of Monte Carlo Localization using laser range finders.gif|thumb|Animation showing the operation of the Monte Carlo Localization for a robot using laser range finders]]&lt;br /&gt;
Monte Carlo Localization is generically known as a particle filter. The key idea of the method is to represent the posterior belief &amp;lt;math&amp;gt;Bel(l)&amp;lt;/math&amp;gt;by a set of &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; weighted, random samples or &#039;&#039;particles&#039;&#039; &amp;lt;math&amp;gt;S = {(s_i | i  = 1...N)}&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. A sample set constitutes a discrete approximation of a probability distribution&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In MCL the samples are of type &amp;lt;math&amp;gt;\langle\langle x,y,\theta\rangle,p\rangle&amp;lt;/math&amp;gt;where &amp;lt;math&amp;gt;x,y,\theta&lt;br /&gt;
&amp;lt;/math&amp;gt;represent the coordinates on the 2D map and the robot orientation, while &amp;lt;math&amp;gt;p \ge 0&amp;lt;/math&amp;gt; represents a discrete probability of the likelihood of the particle&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
MCL operates in two steps, the robot &#039;&#039;motion step&#039;&#039; and the &#039;&#039;sensor reading step&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;When the robot moves&#039;&#039; MCL generates &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;new samples that approximate the robot&#039;s position after the motion command. Each sample is generated by randomly drawing a sample from the previously computed sample set with likelihood determined by their  p-values&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. Let &amp;lt;math&amp;gt;l&#039;&amp;lt;/math&amp;gt; denote the position of this sample. The new sample’s &amp;lt;math&amp;gt;l&amp;lt;/math&amp;gt; is then generated by generating a single, random sample from &amp;lt;math&amp;gt;P(l|l&#039;,a)&amp;lt;/math&amp;gt;, using the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt; as observed. The p-value of the new sample is &amp;lt;math&amp;gt;N^{-1}&amp;lt;/math&amp;gt;.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Sensor readings&#039;&#039; are incorporated by re-weighting the sample set, in a way that implements Bayes rule in Markov localization. More specifically, let &amp;lt;math&amp;gt;\langle l,p \rangle&lt;br /&gt;
&amp;lt;/math&amp;gt;be a sample&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. Then  &amp;lt;math&amp;gt;p \longleftarrow a P(s | l)&lt;br /&gt;
&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;s&lt;br /&gt;
&amp;lt;/math&amp;gt; is the sensor measurement, and &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt; is a normalization constant that enforces the sum of all p-values to 1&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. The incorporation of sensor readings is typically performed in two phases, one in which &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is multiplied by &amp;lt;math&amp;gt;P(s | l)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;, and one in which the various p-values are normalized.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; &lt;br /&gt;
&lt;br /&gt;
In practice the introduction of a small number of uniformly distributed random samples at each step is recommended to enable re-localization in the rare event that the robot looses track of its position.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====  Properties of MCL  ====&lt;br /&gt;
The main property that makes MCL such a powerful algorithm is that it can universally approximate arbitrary probability distributions. The variance of the importance sampler converges to zero at a rate of &amp;lt;math&amp;gt;1/\sqrt{N}&amp;lt;/math&amp;gt;(under MCL conditions). This creates a clear trade-off of accuracy and computational load&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. However, the true advantage lies in the way MCL places computational resources on regions with high likelihood, where things really matter&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; also known as adaptive sampling. More about this is shown in the second paper description below.&lt;br /&gt;
&lt;br /&gt;
Given there is no [https://en.wikipedia.org/wiki/Discretization discretization] of the space or the data, MCL is able to estimate the state of the robot to any numerical accuracy given sampling size.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
MCL is an online algorithm, meaning that it leads itself nicely to an any-time implementation. This means that the algorithm is able to provide an answer at any time, however the quality of the solution increases with time.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Pedestrian Localization for Indoor Environments ===&lt;br /&gt;
The pedestrian localization problem is the second paper of this page, it builds upon the MCL algorithm and adapts it to the localization of pedestrians in an indoor environment. In particular this paper looks at how a foot-mounted inertial unit, a detailed building model, and a particle filter can be combined to provide absolute positioning, despite the presence of drift in the inertial unit and without knowledge of the user’s initial location&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
There are several differences between robot and pedestrian localisation as presented in the paper. Firstly, the robots used in existing literature have been unable to climb stairs. As a result, a 2-dimensional map of the environment has been sufficient. Secondly, the movement of a robot can be actively controlled. This is clearly not possible in a pedestrian localisation system. Thirdly, robot localisation is typically solved knowing both the relative movement of the robot and measurements obtained from a laser range finder. In contrast, only relative movement information  is used (in the form of step events) to solve the pedestrian localisation problem.&lt;br /&gt;
&lt;br /&gt;
==== 2.5D Mapping ====&lt;br /&gt;
[[File:2.5D map lecture room.png|thumb|315x315px|Illustration of a 2.5D map. It illustrates the way stairs and walls are mapped in the environment.]]&lt;br /&gt;
There are many obstacles that limit the possible movement of a pedestrian within a building. In particular walls are impassable obstacles. In order to enforce such constraints it is necessary to have a computer-readable plan of the building. Since it is reasonable to assume that a pedestrian’s foot is constrained to lie on the floor during the stance phase of the gait cycle, a 2.5-dimensional description of the building (in which each object has a vertical position but no depth) is sufficient&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Hence a  map is defined to be a collection of planar floor polygons.Each floor polygon corresponds to a surface in the building on which a pedestrian’s foot may be grounded&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Each edge of a floor polygon is either an impassable wall or a connection to the edge of another polygon. Connected edges must coexist in the (x,y) plane, however they may be separated in the vertical direction to allow the representation of stairs&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Tracking using a foot mounted IMU ====&lt;br /&gt;
Although we&#039;re dealing with tracking people, the MCL is usually applied to robots, where mobile devices typically use inertial sensors, laser range-finders and computer vision to provide accurate localisation without the requirement of fixed infrastructure&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Applying the same systems to people is, however, fraught with difficulties; laser range-finders and cameras are impractical, we lose the ability to control the subject to maximize our chances of precise localisation, and the techniques are usually developed with a single floor in mind&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
One type of sensor which does seem applicable to people tracking is [https://en.wikipedia.org/wiki/Inertial_measurement_unit inertial measurement units (IMUs)]. An IMU contains three orthogonal rate-gyroscopes and accelerometers, which report angular velocity and acceleration respectively. In principle, it is possible to track the orientation of the IMU by integrating the angular velocity signals. This can then be used to resolve the acceleration samples into the global frame of reference, from which acceleration due to gravity is subtracted. The remaining acceleration can then be integrated twice to track the position of the IMU relative to a known starting point and heading&amp;lt;ref&amp;gt;Titterton, D. H., and J. L. Weston. &#039;&#039;Strapdown Inertial Navigation Technology&#039;&#039;. 2nd ed. IEE Radar, Sonar, Navigation, and Avionics Series 17. Stevenage: Institution of Electrical Engineers, 2004.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For foot-mounted IMUs the cubic-in-time drift problem can be reduced by detecting when the foot is in the stationary stance phase (i.e. in contact with the ground) during each gait cycle. Zero velocity updates (ZVUs) can be applied during this phase, in which the known direction of acceleration due to gravity is used to correct tilt errors which have accumulated during the previous step. The application of such constraints replaces the cubic-in-time error growth with an error accumulation that is linear in the number of steps.&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Particle Propagation for pedestrian localization ====&lt;br /&gt;
The particle propagation for pedestrian localization closely resembles the update given by the robot motion in MCL, with an exception for determining which floor polygon to which the propagated particle is constrained. Initially it is assumed that the particle resides in the floor polygon it came from. In order for it to exit the floor polygon the step vector must intersect one of its edges in the (x,y) plane. There are three cases outlined by the authors:&lt;br /&gt;
# No intersection point is found. The particle must still be constrained to the same polygon.&lt;br /&gt;
# The first intersection C is with a wall. In this case the particle’s weight should be set to equal 0 in the correction step, enforcing the constraint that walls are impassable.&lt;br /&gt;
# The first intersection C is with an edge connecting to another polygon. In this case we update the current polygon to the new connected polygon.&lt;br /&gt;
Since a single step may span multiple connections, the intersection test is repeated between the remainder of the updated polygon. This process continues recursively until one of the first two cases applies.&lt;br /&gt;
&lt;br /&gt;
==== Particle Correction for pedestrian localization ====&lt;br /&gt;
The correction step sets the weight &amp;lt;math&amp;gt;w_t&amp;lt;/math&amp;gt;of a propagated particle. This step is used to enforce wall constraints. If a wall is intersected during the propagation step used to generate the state of the particle, then it is assigned a weight wt = 0&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. If a wall is not intersected, the particle is assigned a weight based on the difference between the height change &amp;lt;math&amp;gt;\delta z&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt; of the current step and the difference in height between the start and end floor polygons. The height change according to the map is given by &amp;lt;math&amp;gt;\delta z_{poly} = Height(poly_t) - Height(poly_{t-1})&lt;br /&gt;
&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Here we are using the change in height reported in the current step as a measurement in the Bayesian framework. Particles whose change in height over the step closely matches the change in height reported in the step event are assigned stronger weights&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. This allows localisation to occur quickly when the user climbs or descends stairs.&lt;br /&gt;
[[File:Algorithm update particle propagation and correction for pedestrian localization.png|none|thumb|787x787px|Algorithm containing particle propagation and correction steps for pedestrian localization.                                                                                                                                         Source: Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.”]]&lt;br /&gt;
&lt;br /&gt;
==== Re-sampling for pedestrian localization ====&lt;br /&gt;
The number of particles needed to represent &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt;to a given level of accuracy depends on the complexity of the distribution, which can vary drastically over time. As a result it can be highly inefficient to use a fixed number of particles. This is particularly true for localisation problems, where the number of particles required to track an object after convergence is typically only a small fraction of the number required to adequately describe the distribution in the early stages of localisation&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Kullback-Leibler divergence (KLD) sampling is used in this framework since likelihood-based adaptation is not well suited for problems where &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt; can be a multi-modal distribution, as is often the case during indoor localisation due to symmetry in the environment&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. The idea of KLD-sampling is to generate a number of particles at each step such that the approximation error introduced by using a sample-based representation of &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt;remains below a specified threshold &amp;lt;math&amp;gt;\epsilon&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since the propagation step in this framework uses control information (in the form of step events), the propagated belief is a reasonable estimate of the posterior.&lt;br /&gt;
&lt;br /&gt;
=== Contributions and Comparisons ===&lt;br /&gt;
Taking MCL and adapting it to a pedestrian tracker requires several modifications. MCL was designed for robot localization, robots can have multiple sensors in them to provide information about the environment. Strapping a laser sensor or camera to an individual can be a difficult demand. For such reason the authors resorted in only using an IMU for localization, which makes localization much more difficult. Given humans move by walking in steps, the researchers exploited the IMU data further by making assumptions on the readings, they included ZVU to understand when a step occurred and assumed all footsteps would be in contact with the floor. This enabled the sole use of the IMU for tracking given the aforementioned assumptions and the expansion of the algorithm to a 3 dimensional environment with the 2.5D mapping trick that was show above.&lt;br /&gt;
&lt;br /&gt;
From the pedestrian adaptation of MCL the main contributions to be carried forward are the concept of expanding the map to higher dimensions and the use of assumptions as constraints to reduce the number of sensors or readings needed to make the algorithm work. &lt;br /&gt;
&lt;br /&gt;
The second paper introduces the use of KLD as a method for adaptive sampling, which may be relevant to general applications of MCL.  &lt;br /&gt;
&lt;br /&gt;
It is challenging to compare accuracy in localization for both approaches as environments and data used are very different and the nature of the algorithm in place is stochastic and can be varied in many ways. That said, reported accuracies for experiments for robotics application of MCL fall between 5-10cm with large enough sampling (~1000-10000 samples)&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; while for the pedestrian application they report an accuracy of 0.5m 75% of the time and 0.73m 95% of the time&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. It can be seen that the removal of observatory sensors is quite damaging to accuracy even when employing assumptions to the data being generated, although given the limited sensors and large map environment, the results can still be useful for the task at hand (Robot localization applications tend to require more accurate localization as their navigation often depends on it, human can do it themselves).&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&lt;br /&gt;
&lt;br /&gt;
[2] Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[3] Titterton, D. H., and J. L. Weston. &#039;&#039;Strapdown Inertial Navigation Technology&#039;&#039;. 2nd ed. IEE Radar, Sonar, Navigation, and Avionics Series 17. Stevenage: Institution of Electrical Engineers, 2004.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
	<entry>
		<id>https://wiki.ubc.ca/index.php?title=Course:CPSC522/Monte_Carlo_Localization&amp;diff=582424</id>
		<title>Course:CPSC522/Monte Carlo Localization</title>
		<link rel="alternate" type="text/html" href="https://wiki.ubc.ca/index.php?title=Course:CPSC522/Monte_Carlo_Localization&amp;diff=582424"/>
		<updated>2020-02-10T21:29:08Z</updated>

		<summary type="html">&lt;p&gt;TommasoDAmico: Added sections and rephrased some sentences based on peer critiques&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Monte Carlo Pedestrian Localization for Indoor Environments ==&lt;br /&gt;
This page gives an overview of the Monte Carlo Localization method (MCL)&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; and demonstrates an extension of the technique to enable for pedestrian tracking within a building environment.&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Principal Author: Tommaso D&#039;Amico&lt;br /&gt;
&lt;br /&gt;
Collaborators: -&lt;br /&gt;
&lt;br /&gt;
Papers Discussed:&lt;br /&gt;
&lt;br /&gt;
- Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&lt;br /&gt;
&lt;br /&gt;
- Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
This page gives an overview of the Monte Carlo Localization method (MCL)&amp;lt;ref name=&amp;quot;:0&amp;quot;&amp;gt;Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&amp;lt;/ref&amp;gt; and demonstrates an extension of the technique to enable for pedestrian tracking within a building environment&amp;lt;ref name=&amp;quot;:1&amp;quot;&amp;gt;Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&amp;lt;/ref&amp;gt;. MCL is analyzed as a technique for  pedestrian localization as it allows for global localization with multi modal probability distributions over continuous space and is shown to be robust to highly noisy sensors. The Pedestrian Localization paper demonstrates a technique to extend MCL to a 2.5D map and illustrates the benefits of it for reducing uncertainty.&lt;br /&gt;
&lt;br /&gt;
=== Builds on ===&lt;br /&gt;
The Monte Carlo Localization method is an algorithm that utilizes [[Course:CPSC522/Particle Filtering|particle filtering]] as a means of localization. It builds on the idea of using [[Course:CPSC522/Probability|Bayesian probability]] with [[Course:CPSC522/Markov Chains|Markov Chains]] to develop a [[Course:CPSC522/Probability|probability distribution]] of the likely states of a robot to aid it in its localization. &lt;br /&gt;
&lt;br /&gt;
For the adaptation of MCL to pedestrians in indoors environments the [https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence Kullback-Leibler divergence (KLD)] sampling algorithm is used to dynamically adapt the number of particles sampled.&lt;br /&gt;
&lt;br /&gt;
=== &#039;&#039;&#039;Related Pages&#039;&#039;&#039; ===&lt;br /&gt;
a related combination of Monte Carlo Method with Markov Chains is the [[Course:CPSC522/MCMC|MCMC algorithm]].&lt;br /&gt;
&lt;br /&gt;
== Content ==&lt;br /&gt;
&lt;br /&gt;
=== Monte Carlo Localization (MCL) ===&lt;br /&gt;
The goal of Monte Carlo Localization is to allow a robot to localize itself in an environment given only an odometry system, a sensor to perceive its surroundings and a 2D map of its navigable environment&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In this page we will also explore the use of MCL without the aid of a sensor to perceive the surroundings&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt; and how it is still capable of providing accurate estimates using different aids and assumptions.&lt;br /&gt;
&lt;br /&gt;
==== Global Localization and Position Tracking ====&lt;br /&gt;
Localization for robotics is often broken down in two categories: Global localization and position tracking&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In position tracking the robot knows its initial position and only needs to accommodate for small errors or drift in the odometry system as it moves. Global Localization is a bit more complex as it looks at identifying the robots&#039; position within its global environment given no initial state.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; This is also referred to as the [https://en.wikipedia.org/wiki/Kidnapped_robot_problem hijacked robot problem]. Since the robot needs to localize itself from scratch, this tends to be a very computationally heavy process and is often omitted by most localization algorithms.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====  &#039;&#039;&#039;MCL Algorithm&#039;&#039;&#039; ====&lt;br /&gt;
[[File:Sca80a0 Animation of Monte Carlo Localization using laser range finders.gif|thumb|Animation showing the operation of the Monte Carlo Localization for a robot using laser range finders]]&lt;br /&gt;
Monte Carlo Localization is generically known as a particle filter. The key idea of the method is to represent the posterior belief &amp;lt;math&amp;gt;Bel(l)&amp;lt;/math&amp;gt;by a set of &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; weighted, random samples or &#039;&#039;particles&#039;&#039; &amp;lt;math&amp;gt;S = {(s_i | i  = 1...N)}&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. A sample set constitutes a discrete approximation of a probability distribution&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. In MCL the samples are of type &amp;lt;math&amp;gt;\langle\langle x,y,\theta\rangle,p\rangle&amp;lt;/math&amp;gt;where &amp;lt;math&amp;gt;x,y,\theta&lt;br /&gt;
&amp;lt;/math&amp;gt;represent the coordinates on the 2D map and the robot orientation, while &amp;lt;math&amp;gt;p \ge 0&amp;lt;/math&amp;gt; represents a discrete probability of the likelihood of the particle&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
MCL operates in two steps, the robot &#039;&#039;motion step&#039;&#039; and the &#039;&#039;sensor reading step&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;When the robot moves&#039;&#039; MCL generates &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;new samples that approximate the robot&#039;s position after the motion command. Each sample is generated by randomly drawing a sample from the previously computed sample set with likelihood determined by their  p-values&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. Let &amp;lt;math&amp;gt;l&#039;&amp;lt;/math&amp;gt; denote the position of this sample. The new sample’s &amp;lt;math&amp;gt;l&amp;lt;/math&amp;gt; is then generated by generating a single, random sample from &amp;lt;math&amp;gt;P(l|l&#039;,a)&amp;lt;/math&amp;gt;, using the action &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt; as observed. The p-value of the new sample is &amp;lt;math&amp;gt;N^{-1}&amp;lt;/math&amp;gt;.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Sensor readings&#039;&#039; are incorporated by re-weighting the sample set, in a way that implements Bayes rule in Markov localization. More specifically, let &amp;lt;math&amp;gt;\langle l,p \rangle&lt;br /&gt;
&amp;lt;/math&amp;gt;be a sample&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. Then  &amp;lt;math&amp;gt;p \longleftarrow a P(s | l)&lt;br /&gt;
&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;s&lt;br /&gt;
&amp;lt;/math&amp;gt; is the sensor measurement, and &amp;lt;math&amp;gt;a&lt;br /&gt;
&amp;lt;/math&amp;gt; is a normalization constant that enforces the sum of all p-values to 1&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. The incorporation of sensor readings is typically performed in two phases, one in which &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is multiplied by &amp;lt;math&amp;gt;P(s | l)&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt;, and one in which the various p-values are normalized.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; &lt;br /&gt;
&lt;br /&gt;
In practice the introduction of a small number of uniformly distributed random samples at each step is recommended to enable re-localization in the rare event that the robot looses track of its position.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====  Properties of MCL  ====&lt;br /&gt;
The main property that makes MCL such a powerful algorithm is that it can universally approximate arbitrary probability distributions. The variance of the importance sampler converges to zero at a rate of &amp;lt;math&amp;gt;1/\sqrt{N}&amp;lt;/math&amp;gt;(under MCL conditions). This creates a clear trade-off of accuracy and computational load&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;. However, the true advantage lies in the way MCL places computational resources on regions with high likelihood, where things really matter&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; also known as adaptive sampling. More about this is shown in the second paper description below.&lt;br /&gt;
&lt;br /&gt;
Given there is no [https://en.wikipedia.org/wiki/Discretization discretization] of the space or the data, MCL is able to estimate the state of the robot to any numerical accuracy given sampling size.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
MCL is an online algorithm, meaning that it leads itself nicely to an any-time implementation. This means that the algorithm is able to provide an answer at any time, however the quality of the solution increases with time.&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Pedestrian Localization for Indoor Environments ===&lt;br /&gt;
The pedestrian localization problem is the second paper of this page, it builds upon the MCL algorithm and adapts it to the localization of pedestrians in an indoor environment. In particular this paper looks at how a foot-mounted inertial unit, a detailed building model, and a particle filter can be combined to provide absolute positioning, despite the presence of drift in the inertial unit and without knowledge of the user’s initial location&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
There are several differences between robot and pedestrian localisation as presented in the paper. Firstly, the robots used in existing literature have been unable to climb stairs. As a result, a 2-dimensional map of the environment has been sufficient. Secondly, the movement of a robot can be actively controlled. This is clearly not possible in a pedestrian localisation system. Thirdly, robot localisation is typically solved knowing both the relative movement of the robot and measurements obtained from a laser range finder. In contrast, only relative movement information  is used (in the form of step events) to solve the pedestrian localisation problem.&lt;br /&gt;
&lt;br /&gt;
==== 2.5D Mapping ====&lt;br /&gt;
[[File:2.5D map lecture room.png|thumb|315x315px|Illustration of a 2.5D map. It illustrates the way stairs and walls are mapped in the environment.]]&lt;br /&gt;
There are many obstacles that limit the possible movement of a pedestrian within a building. In particular walls are impassable obstacles. In order to enforce such constraints it is necessary to have a computer-readable plan of the building. Since it is reasonable to assume that a pedestrian’s foot is constrained to lie on the floor during the stance phase of the gait cycle, a 2.5-dimensional description of the building (in which each object has a vertical position but no depth) is sufficient&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Hence a  map is defined to be a collection of planar floor polygons.Each floor polygon corresponds to a surface in the building on which a pedestrian’s foot may be grounded&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Each edge of a floor polygon is either an impassable wall or a connection to the edge of another polygon. Connected edges must coexist in the (x,y) plane, however they may be separated in the vertical direction to allow the representation of stairs&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Tracking using a foot mounted IMU ====&lt;br /&gt;
Although we&#039;re dealing with tracking people, the MCL is usually applied to robots, where mobile devices typically use inertial sensors, laser range-finders and computer vision to provide accurate localisation without the requirement of fixed infrastructure&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. Applying the same systems to people is, however, fraught with difficulties; laser range-finders and cameras are impractical, we lose the ability to control the subject to maximize our chances of precise localisation, and the techniques are usually developed with a single floor in mind&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
One type of sensor which does seem applicable to people tracking is [https://en.wikipedia.org/wiki/Inertial_measurement_unit inertial measurement units (IMUs)]. An IMU contains three orthogonal rate-gyroscopes and accelerometers, which report angular velocity and acceleration respectively. In principle, it is possible to track the orientation of the IMU by integrating the angular velocity signals. This can then be used to resolve the acceleration samples into the global frame of reference, from which acceleration due to gravity is subtracted. The remaining acceleration can then be integrated twice to track the position of the IMU relative to a known starting point and heading&amp;lt;ref&amp;gt;Titterton, D. H., and J. L. Weston. &#039;&#039;Strapdown Inertial Navigation Technology&#039;&#039;. 2nd ed. IEE Radar, Sonar, Navigation, and Avionics Series 17. Stevenage: Institution of Electrical Engineers, 2004.&amp;lt;/ref&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For foot-mounted IMUs the cubic-in-time drift problem can be reduced by detecting when the foot is in the stationary stance phase (i.e. in contact with the ground) during each gait cycle. Zero velocity updates (ZVUs) can be applied during this phase, in which the known direction of acceleration due to gravity is used to correct tilt errors which have accumulated during the previous step. The application of such constraints replaces the cubic-in-time error growth with an error accumulation that is linear in the number of steps.&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Particle Propagation for pedestrian localization ====&lt;br /&gt;
The particle propagation for pedestrian localization closely resembles the update given by the robot motion in MCL, with an exception for determining which floor polygon to which the propagated particle is constrained. Initially it is assumed that the particle resides in the floor polygon it came from. In order for it to exit the floor polygon the step vector must intersect one of its edges in the (x,y) plane. There are three cases outlined by the authors:&lt;br /&gt;
# No intersection point is found. The particle must still be constrained to the same polygon.&lt;br /&gt;
# The first intersection C is with a wall. In this case the particle’s weight should be set to equal 0 in the correction step, enforcing the constraint that walls are impassable.&lt;br /&gt;
# The first intersection C is with an edge connecting to another polygon. In this case we update the current polygon to the new connected polygon.&lt;br /&gt;
Since a single step may span multiple connections, the intersection test is repeated between the remainder of the updated polygon. This process continues recursively until one of the first two cases applies.&lt;br /&gt;
&lt;br /&gt;
==== Particle Correction for pedestrian localization ====&lt;br /&gt;
The correction step sets the weight &amp;lt;math&amp;gt;w_t&amp;lt;/math&amp;gt;of a propagated particle. This step is used to enforce wall constraints. If a wall is intersected during the propagation step used to generate the state of the particle, then it is assigned a weight wt = 0&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. If a wall is not intersected, the particle is assigned a weight based on the difference between the height change &amp;lt;math&amp;gt;\delta z&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/math&amp;gt; of the current step and the difference in height between the start and end floor polygons. The height change according to the map is given by &amp;lt;math&amp;gt;\delta z_{poly} = Height(poly_t) - Height(poly_{t-1})&lt;br /&gt;
&amp;lt;/math&amp;gt;&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Here we are using the change in height reported in the current step as a measurement in the Bayesian framework. Particles whose change in height over the step closely matches the change in height reported in the step event are assigned stronger weights&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. This allows localisation to occur quickly when the user climbs or descends stairs.&lt;br /&gt;
[[File:Algorithm update particle propagation and correction for pedestrian localization.png|none|thumb|787x787px|Algorithm containing particle propagation and correction steps for pedestrian localization.                                                                                                                                         Source: Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.”]]&lt;br /&gt;
&lt;br /&gt;
==== Re-sampling for pedestrian localization ====&lt;br /&gt;
The number of particles needed to represent &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt;to a given level of accuracy depends on the complexity of the distribution, which can vary drastically over time. As a result it can be highly inefficient to use a fixed number of particles. This is particularly true for localisation problems, where the number of particles required to track an object after convergence is typically only a small fraction of the number required to adequately describe the distribution in the early stages of localisation&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Kullback-Leibler divergence (KLD) sampling is used in this framework since likelihood-based adaptation is not well suited for problems where &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt; can be a multi-modal distribution, as is often the case during indoor localisation due to symmetry in the environment&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. The idea of KLD-sampling is to generate a number of particles at each step such that the approximation error introduced by using a sample-based representation of &amp;lt;math&amp;gt;Bel(s)&lt;br /&gt;
&amp;lt;/math&amp;gt;remains below a specified threshold &amp;lt;math&amp;gt;\epsilon&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since the propagation step in this framework uses control information (in the form of step events), the propagated belief is a reasonable estimate of the posterior.&lt;br /&gt;
&lt;br /&gt;
=== Contributions and Comparisons ===&lt;br /&gt;
Taking MCL and adapting it to a pedestrian tracker requires several modifications. MCL was designed for robot localization, robots can have multiple sensors in them to provide information about the environment. Strapping a laser sensor or camera to an individual can be a difficult demand. For such reason the authors resorted in only using an IMU for localization, which makes localization much more difficult. Given humans move by walking in steps, the researchers exploited the IMU data further by making assumptions on the readings, they included ZVU to understand when a step occurred and assumed all footsteps would be in contact with the floor. This enabled the sole use of the IMU for tracking given the aforementioned assumptions and the expansion of the algorithm to a 3 dimensional environment with the 2.5D mapping trick that was show above.&lt;br /&gt;
&lt;br /&gt;
From the pedestrian adaptation of MCL the main contributions to be carried forward are the concept of expanding the map to higher dimensions and the use of assumptions as constraints to reduce the number of sensors or readings needed to make the algorithm work. &lt;br /&gt;
&lt;br /&gt;
The second paper introduces the use of KLD as a method for adaptive sampling, which may be relevant to general applications of MCL.  &lt;br /&gt;
&lt;br /&gt;
It is challenging to compare accuracy in localization for both approaches as environments and data used are very different and the nature of the algorithm in place is stochastic and can be varied in many ways. That said, reported accuracies for experiments for robotics application of MCL fall between 5-10cm with large enough sampling (~1000-10000 samples)&amp;lt;ref name=&amp;quot;:0&amp;quot; /&amp;gt; while for the pedestrian application they report an accuracy of 0.5m 75% of the time and 0.73m 95% of the time&amp;lt;ref name=&amp;quot;:1&amp;quot; /&amp;gt;. It can be seen that the removal of observatory sensors is quite damaging to accuracy even when employing assumptions to the data being generated, although given the limited sensors and large map environment, the results can still be useful for the task at hand (Robot localization applications tend to require more accurate localization as their navigation often depends on it, human can do it themselves).&lt;br /&gt;
&lt;br /&gt;
== Annotated Bibliography ==&lt;br /&gt;
[1] Fox, Dieter, Wolfram Burgard, Frank Dellaert, and Sebastian Thrun. “Monte Carlo Localization: Efﬁcient Position Estimation for Mobile Robots,” In &#039;&#039;Proc&#039;&#039;.&#039;&#039;of the IEEE International Conference on Robotics &amp;amp; Automation (ICRA)&#039;&#039;. 1999, 7.&lt;br /&gt;
&lt;br /&gt;
[2] Woodman, Oliver, and Robert Harle. “Pedestrian Localisation for Indoor Environments.” In &#039;&#039;Proceedings of the 10th International Conference on Ubiquitous Computing - UbiComp ’08&#039;&#039;, 114. Seoul, Korea: ACM Press, 2008. &amp;lt;nowiki&amp;gt;https://doi.org/10.1145/1409635.1409651&amp;lt;/nowiki&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[3] Titterton, D. H., and J. L. Weston. &#039;&#039;Strapdown Inertial Navigation Technology&#039;&#039;. 2nd ed. IEE Radar, Sonar, Navigation, and Avionics Series 17. Stevenage: Institution of Electrical Engineers, 2004.&lt;br /&gt;
&lt;br /&gt;
== To Add ==&lt;br /&gt;
Put links and content here to be added. This does not need to be organized, and will not be graded as part of the page. If you find something that might be useful for a page, feel free to put it here.&lt;/div&gt;</summary>
		<author><name>TommasoDAmico</name></author>
	</entry>
</feed>