From a very young age humans are able to learn how to count, even objects seen for the very first time. However, machine learning models are not able to replicate this success, and most counting pipelines resolve to highly engineered approaches (relying on auxiliary tasks such as classification, identification, or segmentation) which are not only inaccurate, but also have difficulty in counting out-of-distribution objects. In this project, we will address this problem and teach an agent to count items in images in an iterative manner, similarly to how children learn to count.