arXiv 2201.02177
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
By Alethea Power, Yuri Burda, et al.
Published 2022-01-06
Mindmap
Browse the paper's core ideas, clusters, and relationships in a structured outline.
In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can be studied in great detail. In some situations we show that neural networks learn through a process of "grokking" a pattern in the data, improving generalization performance from random chance level t…